Cloudflare Clef and Clef-flash Turn Agent Inputs Into Decisions
marktechpost.com

Cloudflare Clef and Clef-flash Turn Agent Inputs Into Decisions

Tech News
3 min read

Published by AINave Editorial

TL;DRCloudflare’s Clef and Clef-flash answer typed questions about an input state with probabilities, rather than generating free-form text. The distinction could simplify fixed agent decisions, but reported performance varies by task and has not been independently replicated.

Cloudflare has released Clef and Clef-flash, open-weight models designed to return structured decisions rather than chatbot replies. Give either model an input state and typed questions, and it returns probabilities for the permitted answers. That means an agent can receive a score or choice directly instead of generating prose that another system must parse. Both models use the Apache 2.0 license and work with TypeSafe AI’s Jev API, according to the release report.

A fixed answer space changes the output

The reported interface supports three question types: yes-or-no, a choice among named options, and a score against an ordered rubric. A request on Workers AI can include up to 64 questions and four images. The reported input formats also include text, JSON, images and video. That makes the models suited to turning a defined state into machine-readable judgments, rather than handling an open-ended conversation.

The design reflects that narrower job. The backbone makes a single prefill pass over the state and questions, then a joint schema head routes evidence among the fields and scores the options. The probabilities come from those scores. In practical terms, related questions can be evaluated together while the output remains constrained to the schema, as Cloudflare’s described architecture lays out.

Clef is larger; Clef-flash is quicker in one reported test

Clef is a 27B model based on Qwen3.8-27B; Clef-flash is a 9B model based on Qwen3.5-9B. Cloudflare’s internal Decision Index run reported median latencies of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash. Those are measurements from that run, not general latency guarantees. The size and latency figures are reported in the model comparison.

Both models are available through Workers AI, its REST API and AI Gateway, and their weights are on Hugging Face for self-hosting. The report says model cards list testing on a single H200 with BF16 weights; that is a stated test configuration, not a universal hardware requirement. Cloudflare’s announced reinforcement-learning tuning service for private data is also at an early stage: it starts with forward-deployed engineers, while self-serve access is a later plan.

Benchmark leads depend on the task

On Cloudflare’s 10-benchmark shortlist from Decision Index 0.2.1, one of the Clef models ranked first on seven tests. The report gives examples where Clef led Jev, including BANKING77 at 94.20 versus 79.74 and CLINC150+OOS at 97.43 versus 89.27. But Jev led on GPQA Diamond, MMLU-Pro and BBH, so the results do not point to a universal winner. The figures are vendor-reported and have not been independently replicated, as the benchmark summary notes.

That mix is the useful distinction: Clef is a specialized decision interface, not a general replacement for a language model. Its constrained outputs may remove a parsing step in workflows with known questions, while the choice between the two sizes involves a trade-off in the latency figures Cloudflare reported. Whether that trade-off holds for a particular deployment remains a workload-specific question.

FAQs

They are open-weight decision models that return probabilities for answers to typed questions about an input state. Clef is 27B and Clef-flash is 9B; both use Apache 2.0, according to the release report.

Sources

Latest Tech News