ClefBenchClefBench
  • Playground
  • Our results
  • User guide
  • API pricing
  • Pricing
  • Field notes
ClefBenchClefBench

OUR RESULTS / 2026-10-08

Clef, Clef-flash and Jev benchmarks.

We asked Clef, Clef-flash and Jev to sort customer messages, classify invoices and rate risk. Here is what happened.

60

invented sample tasks

540

completed model runs

0

failed attempts

Answers and waiting time

ModelYes/no correctChoice correctAverage rating differenceTypical wait95% finished within
Clef98.33%96.67%0.119 levels0.64 seconds0.94 seconds
Clef-flash96.67%96.67%0.197 levels0.44 seconds0.51 seconds
Jev 1.1396.67%100.00%0.099 levels0.36 seconds0.43 seconds

Higher correct-answer percentages are better. A smaller rating difference means the answer was closer to our expected level. Waiting times include our internet connection and can differ from yours.

Clef had the highest yes/no accuracy in these examples. Jev had the highest choice accuracy and the smallest rating difference. These results do not establish a winner for every task.

How we compared them

  • We wrote 20 customer-message examples, 20 invoice examples and 20 risk examples.
  • Each task had a yes/no question, a choice and a rating from 0 to 3. We set the expected answers before asking any model.
  • Every model answered each task three times, for 180 attempts per model. Repeating a task does not create a new independent example.
  • A yes answer needed at least 50% likelihood. Choices had to match the expected answer. Ratings were compared with the expected level.
  • Failed attempts would count as incorrect. All attempts in this comparison finished successfully. We requested a fresh answer each time; we cannot see whether a model provider reused earlier work.

What to keep in mind

These are short, invented English examples. They do not represent every customer message, invoice or difficult decision. We wrote the expected answers ourselves; they have not had an independent second review.

Perplexity and Strands were not tested here. Try examples from your own work before choosing a model. A model’s confidence is not a guarantee that its answer is correct.

The model providers reported a total charge of $0.027998 for these tests, before account funding fees and taxes.

A closer look at each kind of task

Clef

support

Yes/no correct: 100.0%
Choice correct: 100.0%
Average rating difference: 0.148 levels

invoice

Yes/no correct: 100.0%
Choice correct: 100.0%
Average rating difference: 0.072 levels

risk

Yes/no correct: 95.0%
Choice correct: 90.0%
Average rating difference: 0.135 levels

Clef-flash

support

Yes/no correct: 95.0%
Choice correct: 95.0%
Average rating difference: 0.157 levels

invoice

Yes/no correct: 100.0%
Choice correct: 100.0%
Average rating difference: 0.200 levels

risk

Yes/no correct: 95.0%
Choice correct: 95.0%
Average rating difference: 0.232 levels

Jev 1.13

support

Yes/no correct: 100.0%
Choice correct: 100.0%
Average rating difference: 0.053 levels

invoice

Yes/no correct: 100.0%
Choice correct: 100.0%
Average rating difference: 0.148 levels

risk

Yes/no correct: 90.0%
Choice correct: 100.0%
Average rating difference: 0.094 levels

Explore the evidence

Download the tasks, expected answers and complete results. You can try the same tasks in the playground and compare them with our recorded answers.

Sample tasks and expected answers ↗Original test plan ↗Complete answers ↗Results overview ↗

Completed on 2026-10-08. These original results remain available when later comparisons are added.

Try your own task ↗See plans and costs ↗
ClefBenchClefBench

ClefBench is an independent service. Not affiliated with Cloudflare, TypeSafe, Perplexity or Strands.

Explore
  • Playground
  • Our results
  • Cloudflare Clef guide
Learn
  • User guide
  • Pricing
  • Field notes
Compare
  • API pricing
  • Clef vs Jev
  • Clef vs Perplexity
  • Clef vs Strands
Service
  • Privacy
  • Terms & refunds
  • Cookies
  • Contact
© 2026 ClefBench. All Rights Reserved.