ClefBenchClefBench
  • Playground
  • Our results
  • User guide
  • API pricing
  • Pricing
  • Field notes
ClefBenchClefBench
Decision confidence: thresholds and human review
2026/10/08 · Updated 2026/10/10

Decision confidence: thresholds and human review

Build a review policy around labeled cases, error costs and model disagreement. See why probability and confidence are different fields.

A model's confidence is useful only after you understand the field and test it against your own expected outcomes. A high value is not a guarantee that a decision is correct, and values from different providers are not automatically comparable.

Read the right field

FieldMeaningCommon mistake
noulProbability of yes for the exact binary questionTreating it as general urgency or overall answer quality
choice.probabilitiesDistribution over the allowed choicesIgnoring which options were supplied
choice.confidenceProvider-defined certainty measureAssuming it equals the highest option probability
scoreWeighted index on the supplied ordered scaleReading a 0–3 severity value as a percentage

TypeSafe documents a choice-confidence calculation relative to the number of choices. For example, its formula gives confidence 0.4 when the maximum probability is 0.6 across three choices. That is a provider-specific example, not a conversion rule for every model. TypeSafe confidence reference

Perplexity also returns confidence separately from option probabilities. Preserve those fields when analyzing results rather than replacing one with the other. Perplexity question types

Inspect a disagreement before choosing a threshold

Recorded sample support-06, repetition 1, 2026-10-08. This is a saved synthetic example, not a new model call.

Our entire finance team cannot open invoices today. Other services work.

Frozen reference: service-wide incident = false; responsible team = billing; severity = 2 on the 0–3 scale.

Complete recorded input (JSON)
{
  "state": "Our entire finance team cannot open invoices today. Other services work.",
  "questions": {
    "urgent": {
      "type": "noul",
      "instructions": "Is this a service-wide incident affecting all customers or all tenants? A single company or user does not count."
    },
    "category": {
      "type": "choice",
      "instructions": "Select the primary responsible team.",
      "criteria": {
        "technical": "Software, availability, access or documentation issue",
        "billing": "Charges, invoices, receipts or payments",
        "sales": "Quotes, pricing before purchase or demonstrations"
      }
    },
    "severity": {
      "type": "score",
      "instructions": "Rate the stated impact, using the highest applicable level.",
      "criteria": [
        "Information request, cosmetic documentation issue, or no impact",
        "One person affected or a minor issue with a workaround",
        "An entire company or department affected, not all customers",
        "All customers affected by an outage, lost access, or systemic erroneous charges"
      ]
    }
  }
}
ModelP(service-wide)TeamP(chosen team)Choice confidenceSeverityTime
Clef0.4724billing0.76780.45461.9968759 ms
Clef-flash0.2207technical0.80960.5321.9967463 ms
Jev 1.130.15billing0.80.72.02420 ms

Complete sample set · Original answers · Method and limitations

The binary question in this record asks about a service-wide incident, not whether the finance team needs a prompt reply. All three probabilities fall below the evaluation's 0.5 threshold, while the choice answers disagree about billing versus technical responsibility.

Our frozen reference assigns billing. “Cannot open invoices” can also suggest a technical fault, so the disagreement exposes a category boundary worth reviewing. A threshold cannot repair an ambiguous definition of the right answer.

Draft a policy, then validate it

The following is an illustrative policy for a low-impact classification workflow, not a tested production threshold:

ConditionCandidate action
Required context missing or option definitions overlapRequest information or review manually
Material customer impact, model disagreement or unsupported outputRoute to a person
Binary probability between 0.10 and 0.90Review rather than auto-accept
Binary probability outside that intervalConsider automation only after validation on held-out labeled cases

Those numbers are deliberately example settings. The right policy depends on whether a missed positive or a false alarm is more costly. Do not transfer a binary threshold onto a provider's confidence field or a score output.

Measure the policy on cases it has not seen

  1. Write a labeling rubric and have ambiguous labels reviewed before evaluation.
  2. Split representative cases into a set for choosing the threshold and a separate held-out set for checking it.
  3. Record the proportion automatically handled, errors among those decisions, missed important cases and the human-review workload.
  4. Examine costly mistakes individually. Revisit the question or options when the label itself is unclear.
  5. Recheck after changes to the model version, input distribution or provider route.

Do not report the percentage of accepted decisions as accuracy. A policy that reviews almost everything may have few automatic errors while saving little work. Keep both coverage and error counts visible.

Try it in ClefBench

The result's yes-probability threshold is a local reading aid: changing it does not rerun inference or trigger a downstream action. Use the playground to inspect distributions and the user guide to understand controls. The public benchmark is a small synthetic dataset, not evidence that the illustrative policy above is calibrated for your production tasks.

Provider semantics checked October 10, 2026. Recorded example: original ClefBench evaluation of October 8, 2026.

All Posts

Author

avatar for ClefBench
ClefBench

Categories

  • Decision model guides
Read the right fieldInspect a disagreement before choosing a thresholdDraft a policy, then validate itMeasure the policy on cases it has not seenTry it in ClefBench

More Posts

Email triage with Clef and Jev: a worked example
Decision model guides

Email triage with Clef and Jev: a worked example

Inspect a recorded customer-message sample, exact questions and three model answers, then build a routing rubric for your own inbox.

avatar for ClefBench
ClefBench
2026/10/08
Jev alternatives: hosted models and self-hosted Strands
Decision model guides

Jev alternatives: hosted models and self-hosted Strands

Separate hosted Clef and Perplexity APIs from a deployable Strands checkpoint, then compare control, setup work and total cost.

avatar for ClefBench
ClefBench
2026/10/08
ClefBenchClefBench

ClefBench is an independent service. Not affiliated with Cloudflare, TypeSafe, Perplexity or Strands.

Explore
  • Playground
  • Our results
  • Cloudflare Clef guide
Learn
  • User guide
  • Pricing
  • Field notes
Compare
  • API pricing
  • Clef vs Jev
  • Clef vs Perplexity
  • Clef vs Strands
Service
  • Privacy
  • Terms & refunds
  • Cookies
  • Contact
© 2026 ClefBench. All Rights Reserved.