CLEFBENCH / RESEARCH NOTE / 2026-10-10 / UPDATED 2026-10-10
Cloudflare Clef alternatives: five decision models compared
Compare five decision models: three benchmarked, two guide-only. Check licences, self-hosting, context limits, measured accuracy and official input-token costs.
Which decision model should you use?
Five models, five reasons to investigate one. Start with your constraint, then check the evidence and limits below. Clef, Clef-flash and Jev have recorded ClefBench measurements. Perplexity and Strands are guide only — not tested here.
Highest yes/no accuracy in our set; video-capable weights
Clef$0.240/M input · 638 ms median · 98.33% yes/no · 96.67% choice
Lowest current direct API price among our tested models
Clef-flash$0.038/M input · 436 ms median · 96.67% yes/no · 96.67% choice
Text only; exact category choices and shortest wait in our set
Jev 1.13$0.042/M input · 356.5 ms median · 96.67% yes/no · 100.00% choice
Very long states or more than 64 questions
Perplexity Decider v1.1$0.020/M input · Guide only — not tested here
Operate a smaller open model yourself
Strands Decider v19Self-hosted; budget for hardware and operations · Guide only — not tested here
Cloudflare also lists Clef-omni at $0.150 per million input tokens in its official pricing, checked October 10, 2026. It is a pricing-only reference here: we have not tested it or verified its full specification, so it is outside the five-model comparison.
The five models, field by field
Use the licence and self-hosting columns to narrow deployment choices, then inspect measured waiting time and accuracy. Context and question limits refer to the named providers' interfaces, not this website's playground allowances. A larger context window does not establish better accuracy on long documents.
Official specifications and direct prices checked 2026-10-10. Measurements: 2026-10-08, 60 cases × 3 repetitions per model. Licence and price links lead to their sources.
Scroll horizontally for input, price and measured results. Model names stay visible.
| Model / version | Provider | Licence | Self-host | Size | Context | Questions | Input | Official price / 1M input | Our median wait | Our accuracy | Pick it when |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Clef | Cloudflare | Apache 2.0 | Yes | 27B | 65,536 tokens | 1–64 | Text · JSON · image · video¹ | $0.240 | 638 ms (p95 935) | 98.33% yes/no · 96.67% choice | You need yes/no accuracy or video input; verify your hosting route. |
| Clef-flash | Cloudflare | Apache 2.0 | Yes | 9B | 24,576 tokens | 1–64 | Text · JSON · image · video¹ | $0.038 | 436 ms (p95 513) | 96.67% yes/no · 96.67% choice | Cost leads and your state, questions and media fit in 24,576 tokens. |
| Jev 1.13 | TypeSafe AI | Proprietary | No public weights | Not disclosed | 64k per request; 32k for state + longest question | Subject to both token budgets; see provider docs | Text · JSON (no media) | $0.042 | 356.5 ms (p95 431) | 96.67% yes/no · 100.00% choice | Exact categories matter; Jev led on choices and rating error here. |
| Perplexity Decider v1.1 | Perplexity | Apache 2.0 | Yes; follow checkpoint access and serving instructions | 27B | 262,144-token limit (input must be below it) | 1–128 | Text · JSON · image | $0.020 | Guide only — not tested here | Guide only — not tested here | Investigate its larger documented input and question budgets. |
| Strands Decider v19 | AWS / Strands Labs | Apache 2.0 | Yes | 2B | Runtime-dependent; not verified for v19 here | Runtime-dependent; not verified for v19 here | Text; image support depends on runtime | Self-hosted; no hosted token price | Guide only — not tested here | Guide only — not tested here | You can provision and maintain inference; compute still costs money. |
Clef
- Provider
- Cloudflare
- Licence
- Apache 2.0
- Self-host
- Yes
- Size
- 27B
- Context
- 65,536 tokens
- Questions
- 1–64
- Input
- Text · JSON · image · video¹
- Official price / 1M input
- $0.240
- Our median wait
- 638 ms (p95 935)
- Our accuracy
- 98.33% yes/no · 96.67% choice
- Pick it when
- You need yes/no accuracy or video input; verify your hosting route.
Clef-flash
- Provider
- Cloudflare
- Licence
- Apache 2.0
- Self-host
- Yes
- Size
- 9B
- Context
- 24,576 tokens
- Questions
- 1–64
- Input
- Text · JSON · image · video¹
- Official price / 1M input
- $0.038
- Our median wait
- 436 ms (p95 513)
- Our accuracy
- 96.67% yes/no · 96.67% choice
- Pick it when
- Cost leads and your state, questions and media fit in 24,576 tokens.
Jev 1.13
- Provider
- TypeSafe AI
- Licence
- Proprietary
- Self-host
- No public weights
- Size
- Not disclosed
- Context
- 64k per request; 32k for state + longest question
- Questions
- Subject to both token budgets; see provider docs
- Input
- Text · JSON (no media)
- Official price / 1M input
- $0.042
- Our median wait
- 356.5 ms (p95 431)
- Our accuracy
- 96.67% yes/no · 100.00% choice
- Pick it when
- Exact categories matter; Jev led on choices and rating error here.
Perplexity Decider v1.1
- Provider
- Perplexity
- Licence
- Apache 2.0
- Self-host
- Yes; follow checkpoint access and serving instructions
- Size
- 27B
- Context
- 262,144-token limit (input must be below it)
- Questions
- 1–128
- Input
- Text · JSON · image
- Official price / 1M input
- $0.020
- Our median wait
- Guide only — not tested here
- Our accuracy
- Guide only — not tested here
- Pick it when
- Investigate its larger documented input and question budgets.
Strands Decider v19
- Provider
- AWS / Strands Labs
- Licence
- Apache 2.0
- Self-host
- Yes
- Size
- 2B
- Context
- Runtime-dependent; not verified for v19 here
- Questions
- Runtime-dependent; not verified for v19 here
- Input
- Text; image support depends on runtime
- Official price / 1M input
- Self-hosted; no hosted token price
- Our median wait
- Guide only — not tested here
- Our accuracy
- Guide only — not tested here
- Pick it when
- You can provision and maintain inference; compute still costs money.
¹ Video is a published model capability. Workers AI’s request schema currently exposes images, not a video field. Hosting limits differ; ClefBench accepts text and JSON only. See the full Clef spec.
Jev limits: TypeSafe model reference. Perplexity limits: Decisions API. Strands v19 is our fixed comparison target. The official project has since published v21 and newer checkpoints; check the current reference.
Which of these can you self-host?
Running downloaded weights and reselling a provider's hosted access are different arrangements. The following list adds Laya and Kev as smaller deployment options; neither has been benchmarked by ClefBench.
| Model / weights and licence source | Licence | Self-host | Size | Deployment note |
|---|---|---|---|---|
| Clef | Apache 2.0 | Yes | 27B | Published model code supports images and video |
| Clef-flash | Apache 2.0 | Yes | 9B | Workers AI lists a smaller 24,576-token context |
| Strands Decider v19 | Apache 2.0 | Yes | 2B | Pin the checkpoint and compatible runtime; guide only |
| Perplexity Decider v1.1 | Apache 2.0 | Yes | 27B | Follow checkpoint access and native serving instructions; guide only |
| Laya | Apache 2.0 | Yes | 421M English; 322M multilingual | Small encoders with much shorter input budgets |
| Kev | Apache 2.0 | Yes | 0.8B / 4B / 9B / 27B | Check each variant's tested hardware and context |
| Jev 1.13 | Proprietary | No public weights | Not disclosed | Hosted text/JSON API |
Cloudflare's model card includes Transformers inference code and an SGLang serving recipe using Docker. Hugging Face also exposes integration entry points such as vLLM, but a generic chat-serving example is not proof that it runs the custom decision head correctly. Follow model-specific instructions and validate typed outputs.
Size the hardware for your workload. Cloudflare's published inference example uses a single H200; that is a documented setup, not a universal minimum requirement. A 421M Laya checkpoint and a 9B Clef-flash occupy different memory brackets from Clef 27B. Quantization, context length and concurrency change the requirement. We do not promise that any consumer GPU can run a given variant.
The Apache 2.0 licence permits commercial use subject to its conditions, including applicable notices and redistribution requirements. Operating those weights yourself does not grant permission to resell another company's API. For example, TypeSafe's customer agreement, sections 2.1–2.3, distinguishes integration in customer applications from making the service available as a standalone service. Check the terms for the actual arrangement you plan to use.
Official figures vs what we measured
Cloudflare's published comparison gives median latency of 209.3 ms for Clef, 38.8 ms for Clef-flash and 524.1 ms for Jev 1.13. On that benchmark, Clef's median is about 2.5 times lower than Jev's.
On our recorded path, the order reverses:
| Model | Yes/no accuracy | Choice accuracy | Rating MAE | Median / p95 | 180-run usage cost |
|---|---|---|---|---|---|
| Clef | 98.33% | 96.67% | 0.1185 | 638 / 935 ms | $0.017790 |
| Clef-flash | 96.67% | 96.67% | 0.1965 | 436 / 513 ms | $0.006671 |
| Jev 1.13 | 96.67% | 100.00% | 0.0986 | 356.5 / 431 ms | $0.003537 |
Our times include the network round trip from a local macOS runner through OpenRouter to the pinned provider. They measure end-to-end waiting on that path, not internal model latency. Different tasks, hosting paths and timing boundaries make the two sets unsuitable for a direct speed-ratio claim.
Clef led on yes/no accuracy in our set. Jev led on exact choices, mean absolute rating error and median wait. All three ran the same 60 short synthetic cases, three times each; the reference answers were written by ClefBench and have not received an independent second review. The historical costs above are recorded usage costs, not a recalculation at today's direct API prices.
Read the full methodology and original data, or inspect support routing, invoice review and risk triage. Which errors matter depends on your workflow; these results do not establish a universal winner.
What it costs at your volume
For the hosted APIs listed below, input tokens drive model usage charges. Compare the same workload: 100,000 decisions × 500 input tokens = 50 million input tokens per model. The example includes context and questions; it excludes free allowances, base plans, account fees and taxes. It does not estimate self-hosting costs.
| Model / source | USD / 1M input tokens | Example monthly usage | Price relative to Clef |
|---|---|---|---|
| Clef | $0.240 | $12.00 | Baseline |
| Clef-flash | $0.038 | $1.90 | 15.8% |
| Clef-omni | $0.150 | $7.50 | 62.5% |
| Jev 1.13 | $0.042 | $2.10 | 17.5% |
| Perplexity Decider v1.1 | $0.020 | $1.00 | 8.3% |
At the prices checked on October 10, Clef-flash's direct input rate is about 9.5% lower than Jev's. Its context is 24,576 tokens, compared with Clef's 65,536. Jev has two separate budgets: 64k tokens for the entire request and 32k for state plus the longest question. Cheap input tokens do not help if required evidence is truncated.
Change the workload below using the same calculator as our official API pricing page. Its word-to-token conversion is an estimate; actual tokenization, question length and media affect the bill. Clef-omni's presence in the cost estimate does not imply that it is available in our playground.
LLM pricing calculator for decision models
Estimate direct provider charges for the same text workload across Clef, Clef-flash, Clef-omni, Jev and Perplexity Decider. These are separate from ClefBench subscriptions and credits; a price listing does not imply playground availability.
Official prices in USD per 1 million input tokens. Estimates use each model provider’s published input-token rate.
Monthly API cost by model
Same USD scale across models · shorter bars cost less
- $12.00
$0.240 / 1M input tokens
- $1.90
Clef-flash
Cloudflare Workers AI$0.038 / 1M input tokens
- $7.50
Clef-omni
Cloudflare Workers AI$0.150 / 1M input tokens
- $2.10
Jev 1.13
TypeSafe$0.042 / 1M input tokens
- $1.00
Perplexity Decider v1.1
Perplexity$0.020 / 1M input tokens
Prices checked 2026-10-10. Provider names link to the official pricing sources.
Cloudflare direct: 10,000 free Neurons per day, shared across Workers AI models. Usage beyond that requires Workers Paid, starting at $5/month. Free allowances and the base plan are excluded from these estimates. Cloudflare billing
A rough text-only estimate using 4 tokens per 3 English words. Actual charges depend on tokenization, language and provider. Taxes, account funding fees, base plans and free credits are excluded. Each row estimates one model call per decision.
A ClefBench subscription pays for this website's online features and credits. Direct provider charges and self-hosting infrastructure are separate costs.
What people building with these models report
These are reports from other people's setups, not ClefBench measurements. Sources were checked on 2026-10-10; dates below distinguish an author's publication date from our access date. Star and vote counts are omitted because they change and do not establish quality.
- Local game-agent comparison — Reddit, accessed 2026-10-10. In “Benchmarking decision models is fun — Clef Q8 vs Jev”, the author reports Clef 27B running locally on an RTX 5090 and winning six duels to Jev 1.13's three. The same author reports only one win in 22 matches against a dedicated reinforcement-learning bot. A reply challenges the single-decision setup and points to Jev's ability to evaluate many questions about one state. This is an anecdotal game test, not evidence of general superiority or equivalent throughput.
- Small Workers AI experiments — DEV, published 2026-10-04. 灯里/iku's hands-on report finds simple routing useful but reports missed details in images and multi-intent requests. The author highlights option IDs as model input, sorted object keys, state-tail truncation in the published code and possible interactions between questions. They explicitly do not establish that Workers AI uses every local-code behavior. The article's old Flash price and context are superseded by the current official specifications above.
- Code and deployment tools — GitHub, accessed 2026-10-10. cobanov/awesome-jev collects source-linked projects, while jaredpalmer/kev publishes weights, serving instructions and model-specific hardware notes. Use these repositories to inspect how a deployment works; inclusion in a resource list is not a benchmark endorsement.
Try the interface, then compare your own rules
For direct Cloudflare integration, follow the Clef API tutorial. It includes the authenticated request, model selector and complete question schema. For the model's capabilities, see the full Clef spec.
- Collect representative examples you are permitted to process, including ambiguous boundaries and costly failures.
- Write reference answers before running models; ask a second person to review disputed labels.
- Keep context, question types and allowed answers fixed. A changed rubric is a different test.
- Inspect false positives, false negatives and variation across repetitions, not just aggregate accuracy.
- Measure waiting time and billed input in your intended hosting environment.
You can try a browser comparison or save a private evaluation set with a subscription. For narrower reading, use Clef vs Jev, Clef vs Perplexity Decider and Clef vs Strands Decider.
What we have not tested
- Perplexity Decider, Strands Decider, Clef-omni, Laya and Kev have no ClefBench measurements here.
- The measured set has 60 synthetic short cases: 20 support, 20 invoice and 20 risk. It is not a production-traffic sample.
- Reference answers have no independent second review. Accuracy reflects those references.
- Strands v19 is a fixed documentation target, not the project's latest model.
- We have not evaluated long-context accuracy, images, video, self-hosted deployments or production concurrency.
- Waiting time includes our network path. Hardware, geography and service conditions may change yours.
Questions we get asked
What is the best Cloudflare Clef alternative?
Start with your constraint. Clef-flash has the lowest current direct input price among the three models we tested. Clef led on yes/no accuracy at 98.33%; Jev 1.13 led on category choices at 100.00%. Perplexity Decider v1.1 documents a 262,144-token limit, but we have not tested it. These are task-specific results, not a universal ranking.
Can I self-host Clef?
Yes. Clef and Clef-flash publish Apache 2.0 weights. Cloudflare supplies inference code and serving instructions; follow the decision-model implementation, including its decision head. Hardware, runtime support, weight precision and context length affect what you can run. We have not validated a self-hosted deployment here.
Is Jev open source?
No. Jev 1.13 is a proprietary hosted model with no public weights. TypeSafe's customer agreement restricts offering its service as a standalone service and grants a non-sublicensable access licence. Third-party open implementations are separate projects; they do not make Jev itself open source.
Which decision models can I run locally?
The self-hosting list on this page includes Apache 2.0 releases of Clef, Clef-flash, Strands Decider, Perplexity Decider, Laya and Kev. Follow each checkpoint's access and runtime instructions. Sizes range from Laya's 421M English checkpoint to 27B models. Compute, memory, electricity and maintenance still cost money.
How much cheaper is Clef-flash than Clef?
At the official direct input prices checked on October 10, 2026, Clef costs about 6.3 times as much: $0.240 versus $0.038 per million tokens. Clef-flash is about 84.2% cheaper. At 50 million input tokens, the usage estimates are $12.00 and $1.90, excluding free allowances, account fees and taxes.
What is Clef-flash's context limit?
Cloudflare Workers AI lists 24,576 tokens for Clef-flash, or 37.5% of Clef's 65,536. Questions and media share this budget; long text state may be truncated. Verify your selected hosting route before moving a long-input workload between the two models.
Is ClefBench affiliated with Cloudflare?
No. ClefBench is independent of Cloudflare. We publish a fixed 60-case synthetic evaluation, its method and original results. We identify models we have not tested and separate vendor figures and community reports from our measurements.
