
Cloudflare Clef and Clef-flash Turn Agent Inputs Into Decisions
Published by AINave Editorial
Cloudflare has released Clef and Clef-flash, open-weight models designed to return structured decisions rather than chatbot replies. Give either model an input state and typed questions, and it returns probabilities for the permitted answers. That means an agent can receive a score or choice directly instead of generating prose that another system must parse. Both models use the Apache 2.0 license and work with TypeSafe AI’s Jev API, according to the release report.
A fixed answer space changes the output
The reported interface supports three question types: yes-or-no, a choice among named options, and a score against an ordered rubric. A request on Workers AI can include up to 64 questions and four images. The reported input formats also include text, JSON, images and video. That makes the models suited to turning a defined state into machine-readable judgments, rather than handling an open-ended conversation.
The design reflects that narrower job. The backbone makes a single prefill pass over the state and questions, then a joint schema head routes evidence among the fields and scores the options. The probabilities come from those scores. In practical terms, related questions can be evaluated together while the output remains constrained to the schema, as Cloudflare’s described architecture lays out.
Clef is larger; Clef-flash is quicker in one reported test
Clef is a 27B model based on Qwen3.8-27B; Clef-flash is a 9B model based on Qwen3.5-9B. Cloudflare’s internal Decision Index run reported median latencies of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash. Those are measurements from that run, not general latency guarantees. The size and latency figures are reported in the model comparison.
Both models are available through Workers AI, its REST API and AI Gateway, and their weights are on Hugging Face for self-hosting. The report says model cards list testing on a single H200 with BF16 weights; that is a stated test configuration, not a universal hardware requirement. Cloudflare’s announced reinforcement-learning tuning service for private data is also at an early stage: it starts with forward-deployed engineers, while self-serve access is a later plan.
Benchmark leads depend on the task
On Cloudflare’s 10-benchmark shortlist from Decision Index 0.2.1, one of the Clef models ranked first on seven tests. The report gives examples where Clef led Jev, including BANKING77 at 94.20 versus 79.74 and CLINC150+OOS at 97.43 versus 89.27. But Jev led on GPQA Diamond, MMLU-Pro and BBH, so the results do not point to a universal winner. The figures are vendor-reported and have not been independently replicated, as the benchmark summary notes.
That mix is the useful distinction: Clef is a specialized decision interface, not a general replacement for a language model. Its constrained outputs may remove a parsing step in workflows with known questions, while the choice between the two sizes involves a trade-off in the latency figures Cloudflare reported. Whether that trade-off holds for a particular deployment remains a workload-specific question.





















