CLEFBENCH / RESEARCH NOTE / 2026-10-08 / UPDATED 2026-10-10
Cloudflare Clef: models, limits and examples
Understand Clef and Clef-flash, inspect a recorded decision example, and separate model capabilities from playground limits.
Cloudflare Clef answers defined questions about supplied context. Instead of asking for an open-ended chat reply, you specify what the model should decide and which answers it may return. Typical tasks include routing a support ticket, classifying a purchase and assigning an incident severity.
Use this guide to understand the models. To run a task immediately, open the Clef playground; for measured accuracy and latency, see our independent benchmarks.
Clef and Clef-flash: what differs?
| Official model property | Clef | Clef-flash |
|---|---|---|
| Workers AI identifier | @cf/cloudflare/clef | @cf/cloudflare/clef-flash |
| Model size | 27B parameters | 9B parameters |
| Listed context window | 65,536 tokens | 65,536 tokens |
| Decision types | Yes/no, choice, ordered score | Yes/no, choice, ordered score |
| Available on ClefBench | Free single-model trial; subscription comparisons | Free single-model trial; subscription comparisons |
Specifications were checked against the Clef and Clef-flash model pages on October 10, 2026. A larger model is not a guarantee of a better decision on every task. Test both against the same expected answers.
Model context and hosted input limits
The listed context window is not a promise that every hosting route processes that much state. Cloudflare documents long-text truncation, and OpenRouter's Clef description reports truncation to roughly the first 2,000 tokens of long text state on Workers AI. We have not independently measured that cutoff.
Place the facts needed for a decision near the beginning, remove irrelevant history and verify behavior before relying on long documents. Characters and tokens are different units; the limits of this website are deliberately smaller.
Cloudflare's direct interface also documents image input. ClefBench currently accepts only text or JSON and sends requests through OpenRouter. Direct Workers AI features and free allowances do not automatically become ClefBench features.
Limits of this playground
- Text or JSON context: up to 4,000 characters.
- One to 10 questions; up to 500 characters per question and its conditions.
- Two to 10 choices or ordered rating levels.
- Free access: one Clef or Clef-flash model, up to 5 attempts per day per internet connection, subject to availability. Sharing is required.
- Multiple models and Jev require a ClefBench subscription. Images and video are not accepted in this playground.
The user guide explains model selection, sharing and result recovery.
Three question types and their answers
| Question type | What you define | What to read in the result |
|---|---|---|
noul | A precise yes/no condition | The probability of yes, not a written explanation |
choice | Named options with clear descriptions | The selected option and its probability distribution |
score | Levels in order, lowest first | A weighted level index; it can fall between levels |
A score of 2.4 on a 0–3 scale is not 80% confidence. A provider's confidence field is also not interchangeable with the probability of the selected option. See the confidence review workflow.
A recorded incident example
This example asks whether an outage affects all customers, which team is responsible and how severe the impact is. Expand the JSON to inspect the exact questions and allowed answers before interpreting the output.
Recorded sample support-01, repetition 1, 2026-10-08. This is a saved synthetic example, not a new model call.
Our production API is down for every customer. No requests succeed.
Frozen reference: service-wide incident = true; responsible team = technical; severity = 3 on the 0–3 scale.
Complete recorded input (JSON)
{
"state": "Our production API is down for every customer. No requests succeed.",
"questions": {
"urgent": {
"type": "noul",
"instructions": "Is this a service-wide incident affecting all customers or all tenants? A single company or user does not count."
},
"category": {
"type": "choice",
"instructions": "Select the primary responsible team.",
"criteria": {
"technical": "Software, availability, access or documentation issue",
"billing": "Charges, invoices, receipts or payments",
"sales": "Quotes, pricing before purchase or demonstrations"
}
},
"severity": {
"type": "score",
"instructions": "Rate the stated impact, using the highest applicable level.",
"criteria": [
"Information request, cosmetic documentation issue, or no impact",
"One person affected or a minor issue with a workaround",
"An entire company or department affected, not all customers",
"All customers affected by an outage, lost access, or systemic erroneous charges"
]
}
}
}| Model | P(service-wide) | Team | P(chosen team) | Choice confidence | Severity | Time |
|---|---|---|---|---|---|---|
| Clef | 0.9931 | technical | 0.9599 | 0.8832 | 2.9465 | 1306 ms |
| Clef-flash | 0.9444 | technical | 0.928 | 0.7958 | 2.8898 | 494 ms |
| Jev 1.13 | 0.95 | technical | 1 | 1 | 3 | 343 ms |
Complete sample set · Original answers · Method and limitations
All three models selected the technical team in this recorded run. This is one synthetic sample, not proof of production reliability. The full evaluation includes invoice and risk cases as well as customer messages.
Official API prices and ClefBench subscriptions
| Model / source | USD / 1M input tokens |
|---|---|
| Clef | $0.240 |
| Clef-flash | $0.090 |
These are direct provider input-token rates. Cloudflare Workers AI has a shared daily allocation of 10,000 Neurons, with paid-plan requirements beyond it; account fees and daily usage patterns affect the bill. That allocation is separate from ClefBench's free trial. The official API calculator states its assumptions and excludes free allowances and base-plan fees.
A ClefBench subscription pays for this site's model access and usage credits. It enables private comparisons by default and access to Jev. It does not purchase a Cloudflare Workers plan.
Choose the next step
- Try the same task in Clef and Clef-flash.
- Compare Clef with Jev using our recorded results.
- Review Perplexity's documented alternative or self-hosted Strands; neither was tested in our benchmark.
ClefBench is independent of Cloudflare. Model information checked October 10, 2026. Sources: Cloudflare model documentation, Workers AI billing and OpenRouter's hosting notes.
