ClefBench user guide: your first comparison
A simple guide to choosing a task, asking questions and reading the answers.
This is a guide to the ClefBench browser interface. It does not document a public ClefBench API.
1. Describe your task
Open the playground and choose a sample, or paste your own text. Keep the important details near the beginning and remove names, passwords and confidential information. You can use up to 4,000 characters.
2. Ask clear questions
Add up to ten questions. Choose the answer style that fits each one:
| Answer style | Example |
|---|---|
| Yes or no | Does this customer need an urgent reply? |
| Pick one | Which team should handle this message? |
| Rate on a scale | How serious is this problem? |
For a choice, list one possible answer per line. For a rating, list the levels from lowest to highest. Be specific about what each level means.
3. Compare the answers
Choose models from the dropdown. Guests and free accounts can run either Clef or Clef-flash, one model at a time. Selecting multiple models or Jev prompts guests to register and free accounts to subscribe. Subscribers can choose any one to three models. Read each answer alongside its alternatives and the time it took.
The yes/no value is a probability of yes; a choice has a selected option and distribution; a score is a weighted index on your ordered scale. A provider confidence field is separate from the maximum option probability. It does not guarantee that an answer is correct. If a decision affects someone’s money, health or safety, have a qualified person review it.
4. Choose how to pay
Try up to five free single-model runs a day per internet connection. Free availability is limited. Recorded examples remain available when the trial is full.
The $9.90 monthly subscription unlocks all models and includes 30,000 credits per paid month. Each run shows estimated credits and a maximum hold before submission. Successful models are charged separately using actual input tokens; unused holds and failed models are returned to the original billing period. Unused monthly credits expire with that period.
Credits cover model usage. At 1,000 input tokens, Clef uses 3 credits, Clef-flash 1, and Jev 1. Each successful model is rounded up to a whole credit, with a minimum of 1 credit per model, and its charge never exceeds its disclosed hold. Your available balance shows whole credits you can spend; earlier fractional charges remain unchanged.
Sharing is required for guests and free accounts: the checkbox is checked and cannot be changed. By running, you agree to submit your task and answers for review. Subscribers have sharing off by default and can enable it. Shared tasks and answers are saved for 30 days; an administrator must approve a successful run before it appears on the homepage. Approved submissions remain until removed. Remove personal and confidential information before sharing.
5. Pick up an interrupted result
If you lose your connection, subscribers can restore a recent result for ten minutes. Restoring does not charge you again. Starting a new comparison is a separate attempt.
Copying inputs and results
Use JSON view to inspect the complete context and question definitions. Copy request exports an OpenRouter request body for each selected model; it does not submit a call. Copy result preserves the model response. If you edit a task after running it, the existing results still belong to the previous input.
The yes-probability threshold changes how you inspect an existing result locally; it does not submit another run or take action on the decision. See the confidence workflow and the recorded routing example.
For integration, follow the selected provider's official API documentation using your own provider credentials. A ClefBench subscription does not supply an external provider API key. The official API price comparison links to those sources.
ClefBench