ClefBenchClefBench
ClefBenchClefBench
HomepageClefBench user guidesClefBench user guide: your first comparisonClef API setup: call Clef-flash with your own key

ClefBench user guide: your first comparison

A simple guide to choosing a task, asking questions and reading the answers.

This is a guide to the ClefBench browser interface. It does not document a public ClefBench API.

1. Describe your task

Open the playground and choose a sample, or paste your own text. Keep the important details near the beginning and remove names, passwords and confidential information. You can use up to 4,000 characters.

2. Ask clear questions

Add up to ten questions. Choose the answer style that fits each one:

Answer styleExample
Yes or noDoes this customer need an urgent reply?
Pick oneWhich team should handle this message?
Rate on a scaleHow serious is this problem?

For a choice, list one possible answer per line. For a rating, list the levels from lowest to highest. Be specific about what each level means.

3. Compare the answers

Choose models from the dropdown. Guests and free accounts can run either Clef or Clef-flash, one model at a time. Selecting multiple models or Jev prompts guests to register and free accounts to subscribe. Subscribers can choose any one to three models. Read each answer alongside its alternatives and the time it took.

The yes/no value is a probability of yes; a choice has a selected option and distribution; a score is a weighted index on your ordered scale. A provider confidence field is separate from the maximum option probability. It does not guarantee that an answer is correct. If a decision affects someone’s money, health or safety, have a qualified person review it.

4. Choose how to pay

Try up to five free single-model runs a day per internet connection. Free availability is limited. Recorded examples remain available when the trial is full.

The $9.90 monthly subscription unlocks all models and includes 30,000 credits per paid month. Each run shows estimated credits and a maximum hold before submission. Successful models are charged separately using actual input tokens; unused holds and failed models are returned to the original billing period. Unused monthly credits expire with that period.

Credits cover model usage. At 1,000 input tokens, Clef uses 3 credits, Clef-flash 1, and Jev 1. Each successful model is rounded up to a whole credit, with a minimum of 1 credit per model, and its charge never exceeds its disclosed hold. Your available balance shows whole credits you can spend; earlier fractional charges remain unchanged.

All runs are private by default. Our inference routing service and the selected model provider still process your input. After a successful run, you may explicitly submit the completed input, questions and results for community review within ten minutes. Unpublished submissions are saved for 30 days; only successful, approved submissions appear publicly. Approved submissions remain until removed. Remove sensitive information before running or sharing.

5. Pick up an interrupted result

If you lose your connection, subscribers can restore a recent result for ten minutes. Restoring does not charge you again. Starting a new comparison is a separate attempt.

Copying inputs and results

Use JSON view to inspect the complete context and question definitions. Copy request exports the selected models' request bodies as a reference; it does not submit a call. For direct integration, adapt the body to the provider's documented schema. Copy result preserves the model response. If you edit a task after running it, the existing results still belong to the previous input.

The yes-probability threshold changes how you inspect an existing result locally; it does not submit another run or take action on the decision. See the confidence workflow and the recorded routing example.

For integration, start with calling Clef using your own Cloudflare token, or follow your selected provider's official API documentation. A ClefBench subscription does not supply an external provider API key. The official API price comparison links to those sources.

Try a comparison · See plans · Contact us

Save a repeatable evaluation

Open Saved evaluations, name your evaluation set and add business samples. Each sample has a stable ID, a context, a question set and expected answers. Save your references before running. A subscription supports up to 20 sets, 20 cases per set and 50 saved runs per set.

Select models and review the maximum batch credits. Start the batch and keep the page open: each case reserves and settles credits separately. Progress and results are saved after each case. You can pause or reopen the set, select an unfinished run and continue; submitted calls are never automatically repeated. Failed models remain visible as errors. A fresh run uses credits again.

Modify rules, save a new version and run again. Choose the earlier run as a baseline to see regressions and improvements. Comparisons match stable case IDs, model IDs and question IDs. Different inputs, expected answers, question types or scales are marked as changed cases rather than regressions. Yes/no uses a 50% threshold, choices match exactly, and scores use the nearest scale level.

Export errors to JSON with their complete samples, expected and actual answers, or export the entire run. Explicitly saved sets and results are private and persist until you delete the set or account. After subscription expiry you can still read, export and delete; editing and new or continued runs require an active subscription. Ordinary Playground use still has only the ten-minute paid result recovery window.

ClefBench user guides

Learn the browser workflow, understand model results and find the evidence behind our comparisons.

Clef API setup: call Clef-flash with your own key

Set up Cloudflare credentials, send Clef and Clef-flash decision requests with curl, and understand hosted access versus local installation.

Table of Contents

1. Describe your task2. Ask clear questions3. Compare the answers4. Choose how to pay5. Pick up an interrupted resultCopying inputs and resultsSave a repeatable evaluation