Cloudflare open-sourced Clef and Clef-flash, two fast decision models integrated with Workers AI and compatible with Jev for structured
Cloudflare recently open-sourced two decision models, Clef and Clef-flash, and has integrated them into Workers AI, where developers can call them directly. The model weights are also available under the Apache 2.0 license.Source
Both models also provide a Jev-compatible API, making migration easier for applications that already use Jev.Source
The positioning of this release is very clear: Clef is not a generative model for producing long-form text, but is focused on AI agent classification, judgment, and next-step action selection. Cloudflare also said that Clef and Clef-flash weights can be obtained from Hugging Face for self-deployment, and that model fine-tuning services are now available. In the future, Cloudflare also plans to let customers complete data preparation, fine-tuning, and redeployment through a self-service workflow.Source
The core difference between a decision model and a general LLM lies in output format and inference flow. A general LLM generates text step by step and is suitable for conversations, summaries, and tool operations; Clef instead reads the input first and then directly scores a predefined set of answers, outputting structured decision results and probabilities for each option, making it better suited for structured judgment.Source
Cloudflare's key design points include multimodal input, a 64K context window, and a typed-questions workflow similar to Jev. According to the official information, Clef can read text, JSON, images, or video, and returns the probability for each allowed option; Clef-flash emphasizes speed and is suitable for latency-sensitive automation workflows.Source
On the model foundation, Cloudflare said Clef and Clef-flash are based on Qwen3.8-27B and Qwen3.5-9B respectively, and were trained further for decision tasks. Its key advantage is that the inference method differs from generative models, because the model does not need to generate a block of text and have the system parse it afterward; instead, it scores candidate answers directly, which can greatly shorten processing time for classification and routing tasks.Source
Cloudflare's published results from 43 benchmarks show median latency of 209.3 milliseconds for Clef, 38.8 milliseconds for Clef-flash, and 524.1 milliseconds for Jev. This means that in agentic workflows requiring rapid judgments, the Clef family can significantly reduce waiting time, but differences in accuracy across tests should not be ignored, because Jev still achieved higher scores in some benchmarks.Source
From an engineering perspective, these models are best embedded in customer support routing, ticket prioritization, risk tagging, workflow decisions, and data review scenarios. Their value lies not in generation capability, but in reliably producing structured results that systems can consume directly, thereby reducing manual intervention.Source
For teams using Cloudflare Workers AI, adoption of Clef is relatively low-friction because it can be called directly through the platform without building a full model inference environment. For systems already integrated with a Jev interface, the compatible API also lowers switching costs.Source
For AI agent developers, this means the decision layer can be split out from large generative models, creating a clearer division of responsibilities. Large language models handle understanding, generation, and tool calling, while decision models handle classification, routing, and next-step selection, reducing uncertainty in the overall workflow.Source
For security and operations teams, these models may be used for alert routing, incident severity assessment, and case escalation rules, but organizations also need to design input formats, answer sets, and human review logic more carefully; otherwise, decision errors will directly affect downstream handling.Source
In addition, Cloudflare has also launched fine-tuning services, showing that it is not only providing the model itself but is also trying to enter enterprise-customized decision workflows more deeply. This will make the model closer to industry-specific scenarios, but it also raises requirements for data quality and labeling consistency.Source
If an organization is considering adopting Clef, the first step should be to clearly define the decision scope and avoid having the model handle problems outside the answer set. Input fields, allowed answers, and failure fallback mechanisms should all be defined in advance to ensure the model only makes judgments within a controllable range.Source
The second step should be to establish human review conditions, especially for high-risk cases, low-confidence outputs, or critical workflow checkpoints, where the model should not be relied on for fully automatic decisions. If the model is uncertain, a path for human handling should remain available to reduce the cost of misjudgment.Source
The third step should be to verify data sources, especially when images, text, and JSON are mixed in the input, to guard against junk data, format pollution, and malicious prompt injection. For customer support, ticketing, and alert systems, it is recommended to standardize fields first and then hand them off to the decision model.Source
The fourth step should be to break model evaluation into three dimensions: latency, accuracy, and failure rate, and not look only at speed. Cloudflare's published results already show that the leader can differ by benchmark, so real-world environments need validation using their own data.Source
The fifth step should be to review the supply chain and deployment path. If Workers AI is used, permissions, logging, and API access policies must be confirmed; if self-hosting from Hugging Face, model weights, version management, and update workflows should also be checked to prevent unauthorized changes from affecting decision quality.Source