ClefBenchClefBench
  • Playground
  • Our results
  • User guide
  • API pricing
  • Pricing
  • Field notes
ClefBenchClefBench

CLEFBENCH / RESEARCH NOTE / 2026-10-08 / UPDATED 2026-10-10

Clef vs Jev: accuracy, latency and cost

Compare Clef, Clef-flash and Jev on our fixed 60-case evaluation, with pricing sources, limitations and a practical selection checklist.

Clef and Jev produce structured decisions from supplied context and questions. On ClefBench, subscribers can run both on the same sample, alongside Clef-flash. Free access covers one Clef or Clef-flash model at a time.

What our fixed evaluation measured

We wrote 60 synthetic English cases: 20 support messages, 20 invoices and 20 risk assessments. Each contained a yes/no question, a choice and an ordered rating. Expected answers were fixed before testing. Every model ran each case three times, for 180 runs per model and 540 model runs overall.

ClefBench measurements, 2026-10-08. 60 synthetic cases, three repetitions per model.
ModelYes/no accuracyChoice accuracyRating MAEMedian / p95180-run usage cost
Clef98.33%96.67%0.1185638 / 935 ms$0.017790
Clef-flash96.67%96.67%0.1965436 / 513 ms$0.006671
Jev 1.1396.67%100.00%0.0986356.5 / 431 ms$0.003537

Rating MAE is the mean absolute difference from our expected level on a 0–3 scale; lower is better. Timing includes the runner's network and the OpenRouter/provider path. The cost column is the total reported model usage for 180 runs, not the subscription price or a rate per million tokens.

In this set, Clef had the highest yes/no accuracy. Jev had the highest choice accuracy, lowest rating error and lowest median request time. Those differences do not establish a universal winner. Repeated runs are not additional independent cases, and our labels did not receive an independent second review.

Read the method, per-task breakdown and downloadable evidence before applying these findings to your workload. The recorded Jev responses identify typesafe/jev-1.13-20260917; the site's request route is typesafe/jev-1.13.

Compare capability and hosting separately

DimensionCloudflare ClefTypeSafe Jev 1.13
Official direct interfaceWorkers AITypeSafe System One
Question typesYes/no, choice, ordered scoreYes/no, choice, ordered score
Official input scopeText, JSON and documented multimodal inputText, JSON and text arrays
Published input limit65,536-token directory context; long-state truncation is documented64k tokens per request; state plus the longest question must fit 32k
ClefBench inputText or JSON under the same site limitsText or JSON under the same site limits
ClefBench accessFree single-model trial or subscriptionSubscription

Do not interpret the context figures as equivalent guarantees. Clef hosting has additional truncation notes; see the Clef guide. These differences come from official documentation, not from testing every capability on this website.

Compare costs for your input size

Official direct input prices, checked 2026-10-10.
Model / sourceUSD / 1M input tokens
Clef$0.240
Clef-flash$0.090
Jev 1.13$0.042

Use the API cost calculator for a workload estimate. Use ClefBench pricing for our subscription, sharing rules and monthly credits. These are different purchases.

When would you choose each?

  • Start with Clef when you want to evaluate the Cloudflare model on your own cases or investigate the yes/no result seen in our small evaluation. Validate that advantage on a held-out sample.
  • Include Jev when choice accuracy, text-only decisions or lower model usage cost matter to your evaluation. The observed speed advantage is specific to our runner and hosting path.
  • Include Clef-flash when comparing the smaller Cloudflare model's trade-offs. Its results belong in the same test, not in a separate test with easier inputs.

Define the cost of a wrong route or missed incident before selecting a model. Test easy, ambiguous and costly cases; keep a human-review route for disagreements. The support-ticket example shows why a disagreement can expose an unclear label rather than a simple model failure.

Run your own comparison · Read the user guide

Official information checked October 10, 2026: Cloudflare, TypeSafe model limits, TypeSafe API. Independent measurements remain the original October 8, 2026 dataset.

ClefBenchClefBench

ClefBench is an independent service. Not affiliated with Cloudflare, TypeSafe, Perplexity or Strands.

Explore
  • Playground
  • Our results
  • Cloudflare Clef guide
Learn
  • User guide
  • Pricing
  • Field notes
Compare
  • API pricing
  • Clef vs Jev
  • Clef vs Perplexity
  • Clef vs Strands
Service
  • Privacy
  • Terms & refunds
  • Cookies
  • Contact
© 2026 ClefBench. All Rights Reserved.