⏱️ Reading time: 9 min
Cloudflare just released two artificial intelligence models that don’t generate text: they decide. Clef and Clef-flash classify, route, and escalate tasks using calculated probabilities, and the company claims they already outperform the model that sparked the decision models trend in recent weeks.
📑 En este artículo
It published them on October 2, 2026 under the Apache 2.0 license on Hugging Face, and made them immediately available on Workers AI, its inference platform. It also launched a reinforcement-learning fine-tuning platform so any team can adapt Clef to its own workflow.
TL;DR
- Cloudflare released Clef and Clef-flash, open-source decision models under Apache 2.0, available on Workers AI.
- Clef leads the Jev Decision Index and beats Jev in 3 of 4 workflows evaluated by Typesafe AI.
- Unlike Jev, Clef processes images thanks to a vision encoder and has a 64k-token context window versus 32k.
- Classifying a domain took 2.2s with Clef and 4.7s with the gpt-oss-120b LLM, which also returned fewer categories.
- It also adds a reinforcement-learning fine-tuning platform to customize Clef for specific use cases.
What Clef Is
Clef is one of the new decision models that Cloudflare trains and releases as open source to produce structured outputs with probabilities, such as classifying a message, an image, or a web domain, without retraining it every time a new category appears. It runs on Workers AI and is compatible with the Jev API.
The name comes from music. A clef is the symbol that fixes which note each line of the staff represents, and Cloudflare uses the same idea for its model: Clef sets the domain of the decision and the actions that follow, while the first two letters also wink at the company’s own name.
What Happened
Cloudflare detailed the launch on its official blog, where it published the weights for Clef and Clef-flash on Hugging Face under the Apache 2.0 license, the same one used by most open infrastructure projects because it allows commercial use without asking permission. That same day, both models became available on Workers AI, Cloudflare’s inference platform, so any developer can test them without deploying their own infrastructure.
The company also announced a reinforcement-learning (RL) fine-tuning platform for Clef, designed so teams can train it on their own data without starting from scratch. The announcement came amid the interest generated weeks earlier by Jev System One, Typesafe AI’s decision model that popularized the term outside academic papers.
The Recent Rise of Decision Models
Binary classifiers have existed for years, but the term “decision models” only spread in recent weeks, when Typesafe AI showed that Jev System One returned typed responses with probabilities for specific decisions, instead of free-form text. The difference from a general-purpose LLM is clear: a decision model answers within a fixed set of categories, with a number indicating how confident it is, without needing to be retrained if a new category appears within that same set. This type of classification model is gaining ground because it’s cheaper to operate than a general LLM for repetitive tasks.
A simple example explains it better than the definition. If a support team receives a message, it can run it through a decision model and ask whether it’s urgent and which team should handle it. The model returns a typed response with probabilities (urgent: 92%, billing team: 78%) that the code can use to route the ticket, trigger an alert, or wait for a person to decide. A human no longer needs to be in the middle of every agentic decision.
How Clef Competes on Accuracy and Latency
Clef has two structural differences from Jev. The first is a vision encoder, which lets it classify images in addition to text, something Jev still doesn’t do. The second is context length: 64,000 tokens versus Jev’s 32,000, which allows feeding in more input information before requesting a classification.
Cloudflare published results from several benchmarks used to measure automated decisions, collected in the official announcement. These are some of the most representative:
| Benchmark | Clef | Clef-flash | Jev |
|---|---|---|---|
| BFCL (case exact) | 98.47 | 98.76 | 95.75 |
| API-Bank (accuracy) | 91.93 | 93.11 | 88.19 |
| BANKING77 (macro-F1) | 94.20 | 90.93 | 79.74 |
| CLINC150+OOS (macro-F1) | 97.43 | 66.77 | 89.27 |
| When2Call (accuracy) | 72.37 | 65.58 | 80.97 |
| PhishNChips (accuracy) | 79.60 | 75.05 | 62.55 |
Clef wins four of the six tests in the table. When2Call is the most notable exception, where Jev scores 80.97 against Clef’s 72.37, confirming that no model dominates every scenario equally.
Latency is where the architectural difference stands out most. Workers AI runs the models on Cloudflare’s edge infrastructure, which cuts round-trip time compared to a service centralized in a single region.
| Model | Median Latency | p95 Latency |
|---|---|---|
| Clef | 209.3 ms | 238.6 ms |
| Clef-flash | 38.8 ms | 122.4 ms |
| Jev | 524.1 ms | 536.0 ms |
Cloudflare also tested the models on full workflows, not just isolated benchmarks. In invoice processing, Clef scored 64.7 points against Jev’s 61.8; in customer service, Clef-flash led with 77 versus 76.3 for Clef and 76.0 for Jev. The only category where Jev won clearly was agent trace observability, with 71.6 points versus 68.5 for Clef and 69.8 for Clef-flash.
💭 Key takeaway: Clef doesn’t win across the board. Jev still leads in When2Call, the BRIGHT benchmark, and agent trace observability, so the choice depends on which part of the workflow matters most in each case.
The real-world case Cloudflare uses as an example is its own threat intelligence team. When a domain is passed to Clef, using the Browser Run tool to render it first, the model returns categories with probabilities: a site might come back with a 95% probability of being fashion-related, 85% e-commerce, and less than 1% phishing. Classifying that full domain, fetching the page, rendering it, and classifying it, took 2.2 seconds with Clef. The same flow with gpt-oss-120b, a general-purpose LLM, took 4.7 seconds and returned only two categories instead of the multiple ones Clef provides.
flowchart TD
A["Input: text or image"] --> B["Clef on Workers AI"]
B --> C["Structured output with probabilities"]
C --> D{"High confidence?"}
D -->|"Yes"| E["Automatic action"]
D -->|"No"| F["Escalates to a human"]
The diagram summarizes the full circuit. The input reaches Clef, the model returns a typed output with probabilities, and the code decides whether to act on its own or pass the case to a person based on the confidence threshold each team defines.
Impact and Analysis
Compatibility with the Jev API is the most practical decision behind this launch. A team that already integrated Jev System One can point its calls to Clef without rewriting routing logic, making Cloudflare a low switching-cost alternative rather than a technical overhaul.
Image support opens up use cases Jev doesn’t cover today, such as visual content moderation, scanned document classification, or screenshot review within a support workflow. Combined with the 64,000-token context window, Clef can take in more context per decision without fragmenting the input across multiple calls.
The reinforcement-learning fine-tuning platform is also Cloudflare’s first step toward charging for customization and not just inference. Clef and Clef-flash are free to download and run anywhere, thanks to the Apache 2.0 license, but training them on proprietary data within Workers AI is a separate product.
What’s Next
Cloudflare didn’t publish a date for opening the RL platform to all Workers AI customers, although the announcement already describes it as available alongside the models. What it did confirm is that its own threat intelligence team keeps using Clef in production to classify domains, and that it will keep publishing results against the Jev Decision Index as it trains new versions.
With the weights open on Hugging Face, the community can now compare Clef against other decision models that already appear on the same benchmark, such as DiffusionGemma, Jev Kev 9B, and Laya, without depending on Cloudflare to publish the number.
Try it yourself: download Clef’s weights from Hugging Face or test it directly on Workers AI with your own Cloudflare account to see how long it takes in your real use case.
Frequently Asked Questions
What are decision models?
They are AI models that return typed responses with probabilities within a fixed set of categories, instead of generating free-form text. They’re used to automate repetitive decisions like routing a ticket or classifying a domain.
Does Clef replace an LLM like GPT or Claude?
No, it covers a different task. An LLM reasons in open-ended language and can generate text or call tools; Clef only classifies within known categories, faster and at a lower cost per call.
Are Clef and Clef-flash free?
The weights are published on Hugging Face under the Apache 2.0 license, so running them locally has no licensing cost. Using them hosted on Workers AI follows that platform’s pricing scheme.
What’s the difference between Clef and Clef-flash?
Clef-flash is the latency-optimized version: it responded with a median of 38.8 ms versus Clef’s 209.3 ms in the tests published by Cloudflare, although it loses accuracy on some benchmarks like CLINC150+OOS.
What is the Jev Decision Index?
It’s the public benchmark that compares decision models, including Jev, Clef, Clef-flash, DiffusionGemma, Jev Kev 9B, and Laya, using the same classification tests and workflows.
Can I train my own version of Clef?
Yes, Cloudflare launched a reinforcement-learning fine-tuning platform alongside the models, designed to adapt Clef to each team’s own categories and data.
References
- Cloudflare Blog: official announcement of Clef, Clef-flash, and the RL fine-tuning platform.
- Apache License 2.0: text of the license under which Cloudflare published the weights for both models.
- Cloudflare Workers AI: documentation for the inference platform where Clef and Clef-flash run.
- Hugging Face: repository where Cloudflare published the open weights for Clef and Clef-flash.
📱 Like this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.
Featured image: Foto de Markus Spiske en Unsplash
Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.
Leave a comment
0 Comments