Cloudflare launched the Clef and Clef-flash AI models, open-source decision models designed to produce classifications, structured outputs, and probabilities that can be used programmatically, rather than generating open-ended text as large language models typically do. The company hosts both models on the Workers AI platform and has also made them available on Hugging Face under the Apache 2.0 license.
At the same time, Cloudflare unveiled a new service for fine-tuning Clef using reinforcement learning. The service initially operates through a team of field deployment engineers, with plans to later evolve into a self-service platform that allows customers to collect data, fine-tune the model, and then redeploy it on Cloudflare’s infrastructure.
What Do Decision Models Do?
Decision models target situations in which a system needs to make a specific choice within a workflow, such as determining whether a support message is urgent, assigning it to the billing or technical team, or estimating the severity of its impact. The models return answers written according to defined types, such as a choice, score, or Boolean value, along with probabilities that an application can use to route or escalate a request or refer it to a human employee.
Cloudflare says it tested Clef with its threat intelligence team to classify web domains. In an experiment involving retrieving, rendering, and classifying a website using Browser Run, Clef took 2.2 seconds, compared with 4.7 seconds for the company’s fastest gpt-oss-120b model in the same workflow. Clef also returned multiple classifications with probabilities, including the likelihood that a domain was a fashion or e-commerce website, and a probability of less than 1% that it was a phishing site.
Technical Differences and Performance
Cloudflare says Clef includes a vision encoder for processing images, while Jev, according to the material, focused on text classification. The model also provides a context window of 64,000 tokens, compared with 32,000 for Jev. Clef and Clef-flash use a non-autoregressive architecture during the decision stage, evaluating the correct schema options in parallel instead of generating intermediary text one token at a time.
In the published test results, Clef achieved 98.47% accuracy on the BFCL benchmark, 91.93% on API-Bank, and 94.20% on BANKING77, while Clef-flash achieved 98.76%, 93.11%, and 90.93% on these benchmarks, respectively. The two models did not outperform in every test; on the When2Call benchmark, Jev achieved 80.97%, compared with 72.37% for Clef and 65.58% for Clef-flash, while Jev also remained ahead on the BRIGHT benchmark.
Customization Through Reinforcement Learning
Cloudflare’s training of Clef is based on Qwen, using Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, and then optimizing them for decision-making tasks and probability calibration. The fine-tuning service targets cases such as support-ticket triage, trust and safety report assessment, and bot classification.
The proposed system combines AI Gateway for capturing request data, Workers AI for generating results, Cloudflare Containers for reinforcement-learning environments, and a trainer for updating the model weights, followed by redeployment on Workers AI. The company explains that customizing the model may increase accuracy in a specific domain, but it may come at the expense of general performance.
Why Does This Launch Matter?
Cloudflare presents decision models as a specialized layer between inputs and agent behavior: a relatively small model quickly returns a structured decision, after which a larger language model can execute the resulting action. This design may reduce response times in support, security, and automated-routing paths, but it does not eliminate the need to test calibration and reliability for each use case, particularly when classification probabilities lead to automated actions or sensitive decisions. The results presented also come from evaluations selected by Cloudflare and do not establish Clef’s superiority over all models or domains.