ClefBenchClefBench
  • Playground
  • Our results
  • User guide
  • API pricing
  • Pricing
  • Field notes
ClefBenchClefBench
Email triage with Clef and Jev: a worked example
2026/10/08 · Updated 2026/10/10

Email triage with Clef and Jev: a worked example

Inspect a recorded customer-message sample, exact questions and three model answers, then build a routing rubric for your own inbox.

Email triage combines several decisions: which team owns a message, whether it describes a widespread incident and what impact level applies. Define those separately so a model cannot hide an incorrect assignment inside a plausible summary.

This worked example uses a synthetic customer message from our fixed benchmark. It is not a private customer email or a new model call.

1. Inspect the input and reference labels

Recorded sample support-06, repetition 1, 2026-10-08. This is a saved synthetic example, not a new model call.

Our entire finance team cannot open invoices today. Other services work.

Frozen reference: service-wide incident = false; responsible team = billing; severity = 2 on the 0–3 scale.

Complete recorded input (JSON)
{
  "state": "Our entire finance team cannot open invoices today. Other services work.",
  "questions": {
    "urgent": {
      "type": "noul",
      "instructions": "Is this a service-wide incident affecting all customers or all tenants? A single company or user does not count."
    },
    "category": {
      "type": "choice",
      "instructions": "Select the primary responsible team.",
      "criteria": {
        "technical": "Software, availability, access or documentation issue",
        "billing": "Charges, invoices, receipts or payments",
        "sales": "Quotes, pricing before purchase or demonstrations"
      }
    },
    "severity": {
      "type": "score",
      "instructions": "Rate the stated impact, using the highest applicable level.",
      "criteria": [
        "Information request, cosmetic documentation issue, or no impact",
        "One person affected or a minor issue with a workaround",
        "An entire company or department affected, not all customers",
        "All customers affected by an outage, lost access, or systemic erroneous charges"
      ]
    }
  }
}
ModelP(service-wide)TeamP(chosen team)Choice confidenceSeverityTime
Clef0.4724billing0.76780.45461.9968759 ms
Clef-flash0.2207technical0.80960.5321.9967463 ms
Jev 1.130.15billing0.80.72.02420 ms

Complete sample set · Original answers · Method and limitations

Expand the recorded JSON for the full question set. It defines billing, technical and sales responsibilities and a four-level impact scale. The field named urgent specifically asks whether all customers or tenants are affected; it does not ask whether one team's problem deserves attention.

To reproduce the input, copy that JSON into the playground's JSON editor. Free runs use one Clef or Clef-flash model and require sharing; comparing multiple models or selecting Jev requires a subscription. New runs may differ from the saved results.

2. Explain the difference before picking a winner

Clef and Jev choose billing; Clef-flash chooses technical. The text mentions invoices, but also an inability to open them. Our prewritten reference is billing, while a real organization might give this incident to technical support.

That distinction matters: a benchmark label is not automatically your company's routing policy. Decide whether the rule follows the affected business object or the apparent failure type. Document the boundary and create new reference labels before evaluating revised instructions. Do not relabel the existing dataset silently or attach its recorded results to a changed question.

The rating outputs sit near level 2, matching the stated impact on an entire team rather than all customers. They are weighted scale values, not confidence percentages. The confidence guide explains how to review distributions and disagreements.

3. Build a small routing test

Include in your labeled setWhat it tests
A duplicate charge with working account accessAn unambiguous billing request
A service-wide outageThe positive boundary of the incident question
A single user with a workaroundThe difference between local and broad impact
An invoice that cannot be openedThe billing/technical responsibility boundary
A request for a quote before purchaseThe sales category

These are suggested case types, not additional measurements. Remove personal details and use material you are allowed to share with the selected model provider.

4. Compare useful outcomes

Give each model exactly the same context, option descriptions and rating levels. Count wrong assignments against your rubric, missed widespread incidents and cases needing human review. Then inspect time and cost. A lower token bill may not compensate for more manual correction.

Our Clef vs Jev comparison summarizes all 60 synthetic cases and three repetitions. This single message cannot establish which model is best for an inbox.

Try the recorded input · Read the user guide · Compare official API costs

Example source: support-06, repetition 1, from the October 8 evaluation records. Article updated October 10, 2026.

All Posts

Author

avatar for ClefBench
ClefBench

Categories

  • Decision model guides
1. Inspect the input and reference labels2. Explain the difference before picking a winner3. Build a small routing test4. Compare useful outcomes

More Posts

Jev alternatives: hosted models and self-hosted Strands
Decision model guides

Jev alternatives: hosted models and self-hosted Strands

Separate hosted Clef and Perplexity APIs from a deployable Strands checkpoint, then compare control, setup work and total cost.

avatar for ClefBench
ClefBench
2026/10/08
Decision confidence: thresholds and human review
Decision model guides

Decision confidence: thresholds and human review

Build a review policy around labeled cases, error costs and model disagreement. See why probability and confidence are different fields.

avatar for ClefBench
ClefBench
2026/10/08
ClefBenchClefBench

ClefBench is an independent service. Not affiliated with Cloudflare, TypeSafe, Perplexity or Strands.

Explore
  • Playground
  • Our results
  • Cloudflare Clef guide
Learn
  • User guide
  • Pricing
  • Field notes
Compare
  • API pricing
  • Clef vs Jev
  • Clef vs Perplexity
  • Clef vs Strands
Service
  • Privacy
  • Terms & refunds
  • Cookies
  • Contact
© 2026 ClefBench. All Rights Reserved.