Cloudflare's Clef arrived on October 1, 2026, alongside Clef-flash: two in-house decision models hosted on Workers AI and released with open weights on Hugging Face under Apache 2.0. The post, by Michelle Chen, Alex Reneau, and Kevin Flansburg, frames the pair as a Jev-API-compatible alternative to Typesafe's model and adds a vision encoder that the text-only rival does not offer today.
A decision model does not write an open answer. It classifies a state — a ticket, a domain, a log excerpt — and returns typed outputs with probabilities so code can route, escalate, or hand the case back to a person. Cloudflare describes Clef as a prefill pass over a frozen Qwen backbone, followed by parallel scoring of the valid schema choices. That decision step does not generate tokens one by one.
What changes in practice
Clef uses Qwen3.8-27B as its backbone; Clef-flash starts from Qwen3.5-9B. Both accept images, run with a 64,000-token window (Cloudflare compares that with Jev's 32,000), and are exposed as @cf/cloudflare/clef and @cf/cloudflare/clef-flash. Weights are at Cloudflare/clef and the flash repo.
On the numbers the company itself published, Clef leads BANKING77 (94.20 macro-F1), CLINC150+OOS (97.43), and Amazon ESCI (57.48). Flash leads BFCL case exact (98.76) and home appliances (97.73). Jev still beats both on When2Call (80.97) and BRIGHT (47.52). On Typesafe's workflow suite, Cloudflare says its models win three of four areas — invoices, customer service, and security incidents — and lose on agent-trace observability. These are the publisher's results; broad independent reproduction is not available yet.
Reported median latency is 209.3 ms for Clef and 38.8 ms for flash, versus 524.1 ms for Jev in the same table. An internal domain-classification example, including page rendering, took 2.2 seconds on Clef and 4.7 seconds on gpt-oss-120b, according to the post.
Fine-tuning and what is not self-serve yet
Alongside the weights, Cloudflare is debuting a reinforcement-learning product to calibrate Clef for internal cases such as support triage, trust and safety, and bots. Fine-tuning starts with forward-deployed engineers and is meant to become a self-serve platform later. The company says it does not read, store, or train on requests and responses unless the customer opts into that fine-tuning.
The move comes weeks after Jev defined the "system one" category: bounded, cheap, stable output so an agent can decide before calling a generative model. With open weights, vision, and a compatible API, Cloudflare is trying to turn that category into edge infrastructure, not only a lab product.
Sources
Transparency: This content was created, edited, or reviewed with the aid of artificial intelligence. Information was cross-checked with public posts on X and sources available on the internet. Check the original sources for the full context.
By GeekikiBot