
Decision confidence: thresholds and human review
Build a review policy around labeled cases, error costs and model disagreement. See why probability and confidence are different fields.
A model's confidence is useful only after you understand the field and test it against your own expected outcomes. A high value is not a guarantee that a decision is correct, and values from different providers are not automatically comparable.
Read the right field
| Field | Meaning | Common mistake |
|---|---|---|
noul | Probability of yes for the exact binary question | Treating it as general urgency or overall answer quality |
choice.probabilities | Distribution over the allowed choices | Ignoring which options were supplied |
choice.confidence | Provider-defined certainty measure | Assuming it equals the highest option probability |
score | Weighted index on the supplied ordered scale | Reading a 0–3 severity value as a percentage |
TypeSafe documents a choice-confidence calculation relative to the number of choices. For example, its formula gives confidence 0.4 when the maximum probability is 0.6 across three choices. That is a provider-specific example, not a conversion rule for every model. TypeSafe confidence reference
Perplexity also returns confidence separately from option probabilities. Preserve those fields when analyzing results rather than replacing one with the other. Perplexity question types
Inspect a disagreement before choosing a threshold
Recorded sample support-06, repetition 1, 2026-10-08. This is a saved synthetic example, not a new model call.
Our entire finance team cannot open invoices today. Other services work.
Frozen reference: service-wide incident = false; responsible team = billing; severity = 2 on the 0–3 scale.
Complete recorded input (JSON)
{
"state": "Our entire finance team cannot open invoices today. Other services work.",
"questions": {
"urgent": {
"type": "noul",
"instructions": "Is this a service-wide incident affecting all customers or all tenants? A single company or user does not count."
},
"category": {
"type": "choice",
"instructions": "Select the primary responsible team.",
"criteria": {
"technical": "Software, availability, access or documentation issue",
"billing": "Charges, invoices, receipts or payments",
"sales": "Quotes, pricing before purchase or demonstrations"
}
},
"severity": {
"type": "score",
"instructions": "Rate the stated impact, using the highest applicable level.",
"criteria": [
"Information request, cosmetic documentation issue, or no impact",
"One person affected or a minor issue with a workaround",
"An entire company or department affected, not all customers",
"All customers affected by an outage, lost access, or systemic erroneous charges"
]
}
}
}| Model | P(service-wide) | Team | P(chosen team) | Choice confidence | Severity | Time |
|---|---|---|---|---|---|---|
| Clef | 0.4724 | billing | 0.7678 | 0.4546 | 1.9968 | 759 ms |
| Clef-flash | 0.2207 | technical | 0.8096 | 0.532 | 1.9967 | 463 ms |
| Jev 1.13 | 0.15 | billing | 0.8 | 0.7 | 2.02 | 420 ms |
Complete sample set · Original answers · Method and limitations
The binary question in this record asks about a service-wide incident, not whether the finance team needs a prompt reply. All three probabilities fall below the evaluation's 0.5 threshold, while the choice answers disagree about billing versus technical responsibility.
Our frozen reference assigns billing. “Cannot open invoices” can also suggest a technical fault, so the disagreement exposes a category boundary worth reviewing. A threshold cannot repair an ambiguous definition of the right answer.
Draft a policy, then validate it
The following is an illustrative policy for a low-impact classification workflow, not a tested production threshold:
| Condition | Candidate action |
|---|---|
| Required context missing or option definitions overlap | Request information or review manually |
| Material customer impact, model disagreement or unsupported output | Route to a person |
| Binary probability between 0.10 and 0.90 | Review rather than auto-accept |
| Binary probability outside that interval | Consider automation only after validation on held-out labeled cases |
Those numbers are deliberately example settings. The right policy depends on whether a missed positive or a false alarm is more costly. Do not transfer a binary threshold onto a provider's confidence field or a score output.
Measure the policy on cases it has not seen
- Write a labeling rubric and have ambiguous labels reviewed before evaluation.
- Split representative cases into a set for choosing the threshold and a separate held-out set for checking it.
- Record the proportion automatically handled, errors among those decisions, missed important cases and the human-review workload.
- Examine costly mistakes individually. Revisit the question or options when the label itself is unclear.
- Recheck after changes to the model version, input distribution or provider route.
Do not report the percentage of accepted decisions as accuracy. A policy that reviews almost everything may have few automatic errors while saving little work. Keep both coverage and error counts visible.
Try it in ClefBench
The result's yes-probability threshold is a local reading aid: changing it does not rerun inference or trigger a downstream action. Use the playground to inspect distributions and the user guide to understand controls. The public benchmark is a small synthetic dataset, not evidence that the illustrative policy above is calibrated for your production tasks.
Provider semantics checked October 10, 2026. Recorded example: original ClefBench evaluation of October 8, 2026.
More Posts

Email triage with Clef and Jev: a worked example
Inspect a recorded customer-message sample, exact questions and three model answers, then build a routing rubric for your own inbox.


Jev alternatives: hosted models and self-hosted Strands
Separate hosted Clef and Perplexity APIs from a deployable Strands checkpoint, then compare control, setup work and total cost.

