OUR RESULTS / 2026-10-08
Clef, Clef-flash and Jev benchmarks.
We asked Clef, Clef-flash and Jev to sort customer messages, classify invoices and rate risk. Here is what happened.
invented sample tasks
completed model runs
failed attempts
Answers and waiting time
| Model | Yes/no correct | Choice correct | Average rating difference | Typical wait | 95% finished within |
|---|---|---|---|---|---|
| Clef | 98.33% | 96.67% | 0.119 levels | 0.64 seconds | 0.94 seconds |
| Clef-flash | 96.67% | 96.67% | 0.197 levels | 0.44 seconds | 0.51 seconds |
| Jev 1.13 | 96.67% | 100.00% | 0.099 levels | 0.36 seconds | 0.43 seconds |
Higher correct-answer percentages are better. A smaller rating difference means the answer was closer to our expected level. Waiting times include our internet connection and can differ from yours.
Clef had the highest yes/no accuracy in these examples. Jev had the highest choice accuracy and the smallest rating difference. These results do not establish a winner for every task.
How we compared them
- We wrote 20 customer-message examples, 20 invoice examples and 20 risk examples.
- Each task had a yes/no question, a choice and a rating from 0 to 3. We set the expected answers before asking any model.
- Every model answered each task three times, for 180 attempts per model. Repeating a task does not create a new independent example.
- A yes answer needed at least 50% likelihood. Choices had to match the expected answer. Ratings were compared with the expected level.
- Failed attempts would count as incorrect. All attempts in this comparison finished successfully. We requested a fresh answer each time; we cannot see whether a model provider reused earlier work.
What to keep in mind
These are short, invented English examples. They do not represent every customer message, invoice or difficult decision. We wrote the expected answers ourselves; they have not had an independent second review.
Perplexity and Strands were not tested here. Try examples from your own work before choosing a model. A model’s confidence is not a guarantee that its answer is correct.
The model providers reported a total charge of $0.027998 for these tests, before account funding fees and taxes.
A closer look at each kind of task
Clef
Yes/no correct: 100.0%
Choice correct: 100.0%
Average rating difference: 0.148 levels
Yes/no correct: 100.0%
Choice correct: 100.0%
Average rating difference: 0.072 levels
Yes/no correct: 95.0%
Choice correct: 90.0%
Average rating difference: 0.135 levels
Clef-flash
Yes/no correct: 95.0%
Choice correct: 95.0%
Average rating difference: 0.157 levels
Yes/no correct: 100.0%
Choice correct: 100.0%
Average rating difference: 0.200 levels
Yes/no correct: 95.0%
Choice correct: 95.0%
Average rating difference: 0.232 levels
Jev 1.13
Yes/no correct: 100.0%
Choice correct: 100.0%
Average rating difference: 0.053 levels
Yes/no correct: 100.0%
Choice correct: 100.0%
Average rating difference: 0.148 levels
Yes/no correct: 90.0%
Choice correct: 100.0%
Average rating difference: 0.094 levels
Explore the evidence
Download the tasks, expected answers and complete results. You can try the same tasks in the playground and compare them with our recorded answers.
Completed on 2026-10-08. These original results remain available when later comparisons are added.
