In the third week of September 2026, three major model launches landed back to back, and all three said the same thing: cheaper, at least as good as before. Each came with a chart, and each chart was built from tests the announcing company chose.

Those charts are not useless, but they do not answer your question. Your question is: which one is better at my work, and what does it cost me? Only your own data can answer that. The good news is that you do not need much infrastructure to find out.

What an evaluation set is

An evaluation set (an eval) is a small collection of inputs you give the model together with the answers those inputs should produce. OpenAI's documentation frames the process in three steps: describe the task as an eval configuration, run the eval, then analyse the results and iterate.

The goal is not to judge the model in general. It is to see whether a change you made, switching models, editing a prompt, lowering the effort setting, broke your work.

Step 1: Turn your success criterion into a number

Anthropic's documentation contrasts good and bad criteria with examples. "Safe outputs" is a bad criterion; "less than 0.1% of outputs out of 10,000 trials flagged for toxicity" is a good one. Likewise, instead of "classify sentiment well", say "an F1 score of at least 0.85".

Writing this down is usually the hardest step, and skipping it makes everything after it meaningless. Turn "the customer reply should be good" into something checkable: "the reply names the product the customer asked about and states the return window correctly".

Step 2: Take your examples from real work

Two of Anthropic's three design principles apply here:

  • Be task-specific: your test cases should mirror your real task distribution. Include the edges: irrelevant input, overly long input, harmful input, ambiguous cases.
  • Prioritise volume over quality: in the documentation's words, more questions with slightly lower-signal automated grading beats fewer questions graded by hand.

In practice, 30 to 50 examples is enough to start. Do not invent them; take them from real records: last month's support tickets, bugs from your own codebase, questions people actually asked. Collect the cases the model got wrong before, in particular; those carry the most information.

Step 3: Automate the grading

A test you read by hand gets abandoned on its third run. Anthropic's documentation lists several automated grading methods:

MethodWhere it works
Exact matchClassification: is the label right or not
Cosine similarityConsistency: similar answers to paraphrased questions
ROUGE-LSummarisation: overlap with a reference summary
Model-graded scale (1-5)Subjective qualities like tone, empathy, professionalism
Model-graded binaryYes/no calls such as "was personal data leaked"

If you grade with a model, two rules matter: use a different model from the one producing the answers, and tie the grading to an explicit rubric. The test of a good rubric is whether two experts would give the same answer the same score.

Step 4: Measure cost too

This week's launches were all about cost, so your eval should track three numbers rather than accuracy alone:

  • Success rate.
  • Tokens per task (input, output and cache reads separately).
  • Time per task.

Vendor charts have shifted the same way, talking about cost per task rather than absolute scores. Running that calculation on your own work is the only way to learn what a claim like "40% cheaper" means for you.

A worked example: support replies

Say a model writes the support replies for an online shop. The evaluation set is built like this:

  1. Examples: 40 real support tickets from the past month. 25 of the common kinds (where is my parcel, how do returns work, size exchanges), 10 hard cases (two orders mixed up, partial refunds), 5 edge cases (irrelevant message, angry customer, missing details).
  2. Expected output: not a hand-written perfect reply, but what a correct reply must contain. For instance: "states the return window as 14 days", "repeats the order number", "does not name the courier".
  3. Grading: those items can be checked with exact match or a simple string check. Tone gets a separate model grader that judges tone only.
  4. Cost: each run records total input and output tokens, which gives cost per task.

A set like this takes an afternoon to build and lasts for years. When a new model appears, the work is running two commands and putting two tables side by side.

Three common mistakes

  • Only including easy examples. Every model passes the easy cases; a set only discriminates where models struggle. Half of it should be hard.
  • Never auditing the grader. Model graders get things wrong too. Read and score ten examples by hand now and then and compare; if you disagree, fix the rubric.
  • Looking at a single number. An overall success rate hides which kind of case broke. Break results down by category: if shipping questions improved while returns got worse, the average will not say so.

Step 5: Read the result correctly

The most common mistake here is treating a small gap between two models as a real difference. Anthropic's paper on a statistical approach to evals offers a few corrections:

  • Report the margin of error. From the standard error of the mean you can give a 95% confidence interval: the mean plus and minus 1.96 standard errors.
  • Adjust for clustered questions. Several questions about the same passage are not independent; the paper notes clustered standard errors can be over three times as large as naive ones.
  • Use paired comparisons. Because you test both models on the same questions, paired analysis removes the noise that comes from question difficulty.

In practice: on a 40-example set, the difference between 72% and 75% is probably noise. Either add examples or do not make that gap the reason for a decision.

Step 6: Run it regularly

An eval's value is in repetition, not in a one-off comparison. Run it at three moments:

  1. When switching models or versions.
  2. When changing the prompt.
  3. When trying a lower effort setting or a smaller model to cut cost.

Store the results with the date and model name. Three months later, when someone says "it used to do this better", you will have a record instead of an impression.

Where to keep the set

An evaluation set is not a document; it is part of the code. Three practical rules help:

  • Keep it in the repository. Examples and expected outputs belong beside the application's source, so that when a prompt changes the set is updated in the same commit and the two never drift apart.
  • Scrub customer data. When you take examples from real records, replace names, addresses, order numbers and phone numbers. The data needs to be realistic, not real; moving personal data into a repository is a risk you do not need.
  • Make running it one command. A set that takes more than two lines to run does not get run. A small script that prints a table beats the most expensive evaluation tool.

On a team you can also run it automatically on every change. But build a set that works by hand first, and add the automation once you trust it.

Where to start

The smallest useful thing you can do today: collect 20 cases from the past month where the model got it wrong, write the correct answer for each, run them on two models and record the success rate and tokens per task. Even that is enough to know what to do the next time an announcement says "50% cheaper".