
Email triage with Clef and Jev: a worked example
Inspect a recorded customer-message sample, exact questions and three model answers, then build a routing rubric for your own inbox.
Email triage combines several decisions: which team owns a message, whether it describes a widespread incident and what impact level applies. Define those separately so a model cannot hide an incorrect assignment inside a plausible summary.
This worked example uses a synthetic customer message from our fixed benchmark. It is not a private customer email or a new model call.
1. Inspect the input and reference labels
Recorded sample support-06, repetition 1, 2026-10-08. This is a saved synthetic example, not a new model call.
Our entire finance team cannot open invoices today. Other services work.
Frozen reference: service-wide incident = false; responsible team = billing; severity = 2 on the 0–3 scale.
Complete recorded input (JSON)
{
"state": "Our entire finance team cannot open invoices today. Other services work.",
"questions": {
"urgent": {
"type": "noul",
"instructions": "Is this a service-wide incident affecting all customers or all tenants? A single company or user does not count."
},
"category": {
"type": "choice",
"instructions": "Select the primary responsible team.",
"criteria": {
"technical": "Software, availability, access or documentation issue",
"billing": "Charges, invoices, receipts or payments",
"sales": "Quotes, pricing before purchase or demonstrations"
}
},
"severity": {
"type": "score",
"instructions": "Rate the stated impact, using the highest applicable level.",
"criteria": [
"Information request, cosmetic documentation issue, or no impact",
"One person affected or a minor issue with a workaround",
"An entire company or department affected, not all customers",
"All customers affected by an outage, lost access, or systemic erroneous charges"
]
}
}
}| Model | P(service-wide) | Team | P(chosen team) | Choice confidence | Severity | Time |
|---|---|---|---|---|---|---|
| Clef | 0.4724 | billing | 0.7678 | 0.4546 | 1.9968 | 759 ms |
| Clef-flash | 0.2207 | technical | 0.8096 | 0.532 | 1.9967 | 463 ms |
| Jev 1.13 | 0.15 | billing | 0.8 | 0.7 | 2.02 | 420 ms |
Complete sample set · Original answers · Method and limitations
Expand the recorded JSON for the full question set. It defines billing, technical and sales responsibilities and a four-level impact scale. The field named urgent specifically asks whether all customers or tenants are affected; it does not ask whether one team's problem deserves attention.
To reproduce the input, copy that JSON into the playground's JSON editor. Free runs use one Clef or Clef-flash model and require sharing; comparing multiple models or selecting Jev requires a subscription. New runs may differ from the saved results.
2. Explain the difference before picking a winner
Clef and Jev choose billing; Clef-flash chooses technical. The text mentions invoices, but also an inability to open them. Our prewritten reference is billing, while a real organization might give this incident to technical support.
That distinction matters: a benchmark label is not automatically your company's routing policy. Decide whether the rule follows the affected business object or the apparent failure type. Document the boundary and create new reference labels before evaluating revised instructions. Do not relabel the existing dataset silently or attach its recorded results to a changed question.
The rating outputs sit near level 2, matching the stated impact on an entire team rather than all customers. They are weighted scale values, not confidence percentages. The confidence guide explains how to review distributions and disagreements.
3. Build a small routing test
| Include in your labeled set | What it tests |
|---|---|
| A duplicate charge with working account access | An unambiguous billing request |
| A service-wide outage | The positive boundary of the incident question |
| A single user with a workaround | The difference between local and broad impact |
| An invoice that cannot be opened | The billing/technical responsibility boundary |
| A request for a quote before purchase | The sales category |
These are suggested case types, not additional measurements. Remove personal details and use material you are allowed to share with the selected model provider.
4. Compare useful outcomes
Give each model exactly the same context, option descriptions and rating levels. Count wrong assignments against your rubric, missed widespread incidents and cases needing human review. Then inspect time and cost. A lower token bill may not compensate for more manual correction.
Our Clef vs Jev comparison summarizes all 60 synthetic cases and three repetitions. This single message cannot establish which model is best for an inbox.
Try the recorded input · Read the user guide · Compare official API costs
Example source: support-06, repetition 1, from the October 8 evaluation records. Article updated October 10, 2026.
More Posts

Jev alternatives: hosted models and self-hosted Strands
Separate hosted Clef and Perplexity APIs from a deployable Strands checkpoint, then compare control, setup work and total cost.


Decision confidence: thresholds and human review
Build a review policy around labeled cases, error costs and model disagreement. See why probability and confidence are different fields.

