CLEFBENCH / RESEARCH NOTE / 2026-10-08 / UPDATED 2026-10-10
Clef vs Jev: accuracy, latency and cost
Compare Clef, Clef-flash and Jev on our fixed 60-case evaluation, with pricing sources, limitations and a practical selection checklist.
Clef and Jev produce structured decisions from supplied context and questions. On ClefBench, subscribers can run both on the same sample, alongside Clef-flash. Free access covers one Clef or Clef-flash model at a time.
What our fixed evaluation measured
We wrote 60 synthetic English cases: 20 support messages, 20 invoices and 20 risk assessments. Each contained a yes/no question, a choice and an ordered rating. Expected answers were fixed before testing. Every model ran each case three times, for 180 runs per model and 540 model runs overall.
| Model | Yes/no accuracy | Choice accuracy | Rating MAE | Median / p95 | 180-run usage cost |
|---|---|---|---|---|---|
| Clef | 98.33% | 96.67% | 0.1185 | 638 / 935 ms | $0.017790 |
| Clef-flash | 96.67% | 96.67% | 0.1965 | 436 / 513 ms | $0.006671 |
| Jev 1.13 | 96.67% | 100.00% | 0.0986 | 356.5 / 431 ms | $0.003537 |
Rating MAE is the mean absolute difference from our expected level on a 0–3 scale; lower is better. Timing includes the runner's network and the OpenRouter/provider path. The cost column is the total reported model usage for 180 runs, not the subscription price or a rate per million tokens.
In this set, Clef had the highest yes/no accuracy. Jev had the highest choice accuracy, lowest rating error and lowest median request time. Those differences do not establish a universal winner. Repeated runs are not additional independent cases, and our labels did not receive an independent second review.
Read the method, per-task breakdown and downloadable evidence before applying these findings to your workload. The recorded Jev responses identify typesafe/jev-1.13-20260917; the site's request route is typesafe/jev-1.13.
Compare capability and hosting separately
| Dimension | Cloudflare Clef | TypeSafe Jev 1.13 |
|---|---|---|
| Official direct interface | Workers AI | TypeSafe System One |
| Question types | Yes/no, choice, ordered score | Yes/no, choice, ordered score |
| Official input scope | Text, JSON and documented multimodal input | Text, JSON and text arrays |
| Published input limit | 65,536-token directory context; long-state truncation is documented | 64k tokens per request; state plus the longest question must fit 32k |
| ClefBench input | Text or JSON under the same site limits | Text or JSON under the same site limits |
| ClefBench access | Free single-model trial or subscription | Subscription |
Do not interpret the context figures as equivalent guarantees. Clef hosting has additional truncation notes; see the Clef guide. These differences come from official documentation, not from testing every capability on this website.
Compare costs for your input size
| Model / source | USD / 1M input tokens |
|---|---|
| Clef | $0.240 |
| Clef-flash | $0.090 |
| Jev 1.13 | $0.042 |
Use the API cost calculator for a workload estimate. Use ClefBench pricing for our subscription, sharing rules and monthly credits. These are different purchases.
When would you choose each?
- Start with Clef when you want to evaluate the Cloudflare model on your own cases or investigate the yes/no result seen in our small evaluation. Validate that advantage on a held-out sample.
- Include Jev when choice accuracy, text-only decisions or lower model usage cost matter to your evaluation. The observed speed advantage is specific to our runner and hosting path.
- Include Clef-flash when comparing the smaller Cloudflare model's trade-offs. Its results belong in the same test, not in a separate test with easier inputs.
Define the cost of a wrong route or missed incident before selecting a model. Test easy, ambiguous and costly cases; keep a human-review route for disagreements. The support-ticket example shows why a disagreement can expose an unclear label rather than a simple model failure.
Run your own comparison · Read the user guide
Official information checked October 10, 2026: Cloudflare, TypeSafe model limits, TypeSafe API. Independent measurements remain the original October 8, 2026 dataset.
