Cloudflare Clef shipped on October 1, 2026, alongside Clef-flash, two decision models trained by the Workers AI team. The weights are on Hugging Face under Apache 2.0, and hosted versions already answer at @cf/cloudflare/clef and @cf/cloudflare/clef-flash.
A decision model does not write free text. It takes a state and a question schema and returns typed answers with probabilities — for example, whether a support ticket is urgent and which team should own it. Cloudflare places Clef in the same family as Typesafe's Jev and says the API is compatible, so a workflow can swap backends without a rewrite.
What changes versus Jev
In the official post, Michelle Chen, Alex Reneau, and Kevin Flansburg list three differences. Clef has a vision encoder and can classify images, while Jev, according to Cloudflare, is still text-only. The context window is 64,000 tokens, against Jev's 32,000. Median latency measured by the company is 209.3 ms for Clef and 38.8 ms for Clef-flash, against 524.1 ms for Jev on the same test set.
The larger model freezes Qwen3.8-27B as its backbone; flash uses Qwen3.5-9B. The decision step is not autoregressive: a prefill pass feeds a two-stage attention router, and valid schema choices are scored in parallel. Training combines label-smoothed cross-entropy, a Brier loss for probability calibration, and a stage called RLCD, reinforcement learning for calibrated decisions.
Numbers Cloudflare publishes
On the decision index used by the company, Clef leads on ToolRet (69.19 nDCG@10), BANKING77 (94.20 macro-F1), and CLINC150+OOS (97.43). Clef-flash leads on BFCL case exact (98.76) and home appliances (97.73). Jev is still ahead on When2Call (80.97 versus 72.37) and on BRIGHT. On Typesafe's own workflows, Cloudflare says Clef wins three of four areas: invoices, customer service, and security incidents. Agent-trace observability stays with Jev.
Internally, the Threat Intelligence team tested domain classification with Browser Run: 2.2 seconds to fetch, render, and classify, versus 4.7 seconds for gpt-oss-120b on the same flow, according to the blog. Those times are Cloudflare's own measurements, not an independent lab result.
Fine-tuning is not self-serve yet
The company also announced a reinforcement-learning product to adapt Clef to internal cases. The first stage is done with forward-deployed engineers; a self-serve platform comes later. The promise not to read, store, or train on requests applies to ordinary inference, with the explicit exception of fine-tuning.
For people building agents, the practical point is elsewhere: cheap, bounded classification on the hot path, and an LLM only when the task needs text or tools. Cloudflare Clef does not replace a frontier model. It occupies the stretch where the answer needs to be a label with a probability, not a paragraph.
Sources
- Cloudflare Blog: Introducing Clef
- Cloudflare changelog: Clef on Workers AI
- Hugging Face: Cloudflare/clef
Transparency: This content was created, edited, or reviewed with the assistance of artificial intelligence. Information was cross-checked with public posts on X and sources available on the internet. Consult the original sources for the full context.
By GeekikiBot