📊 Evaluation & Benchmarking
Choose a decision-relevant metric, compare models on the same cases, and separate real improvements from noise or measurement errors.
On this page
Before you start
Probability & Statistics covers rates, sampling and uncertainty. ML Fundamentals introduces training, validation and test splits. You do not need a leaderboard to begin this lesson.
You will learn to turn a product decision into an evaluation, calculate common metrics, compare two systems on the same examples, and diagnose failures in retrieval, generation or the evaluator itself.
Start with the decision
Suppose a classifier flags transactions for review. Missing fraud costs money, but false alarms also consume reviewer time. Before choosing a metric, define the positive class, target population, unit of evaluation, available action and costs of its errors.
An evaluation is a measurement procedure: examples + system configuration + scoring rule + aggregation + uncertainty. Change any of these and a score can move even when the model weights stay the same.
Work through a confusion matrix
Out of 1,000 transactions, 20 are fraudulent. A model flags 30: 12 are fraud and 18 are legitimate.
| Actual / predicted | Flagged | Not flagged |
|---|---|---|
| Fraud | TP = 12 | FN = 8 |
| Legitimate | FP = 18 | TN = 962 |
- Accuracy =
(TP+TN)/N = 974/1000 = 97.4%. - Precision =
TP/(TP+FP) = 12/30 = 40%: how many alerts are correct. - Recall =
TP/(TP+FN) = 12/20 = 60%: how much actual fraud is caught. - F1 =
2TP/(2TP+FP+FN) = 24/50 = 0.48.
Predicting “legitimate” for everything gets 98% accuracy, yet catches no fraud. If a missed fraud costs $50 and a false alarm costs $2, the model's simplified error cost is 8×50 + 18×2 = $436, compared with 20×50 = $1,000 for that baseline. These are assumed costs for this example, excluding review capacity and other outcomes.
If p is a calibrated fraud probability and these are the only action costs, flag when 2(1−p) < 50p, or p > 2/52 ≈ 0.0385. A threshold is a decision rule; 0.5 is not inherently the right one.
Ranking, probabilities and thresholds answer different questions
ROC-AUC summarizes positive-versus-negative ranking across thresholds; PR curves show the precision/recall tradeoff and depend strongly on prevalence. Neither selects an operating threshold or proves that the model's probabilities are calibrated.
Calibration asks whether, among predictions near 0.7, the event occurs about 70% of the time. Use reliability diagrams and proper scores such as log loss or Brier score. Expected Calibration Error summarizes bin-wise gaps, but changes with the binning, sample size and aggregation. Check decision-relevant slices and uncertainty rather than treating one ECE value as certification.
Define empty-denominator handling. Precision for a model issuing zero positive predictions is undefined mathematically; libraries may substitute zero or another convention. Record the convention instead of hiding it.
Compare the same cases
Two assistants answer the same 200 independent cases:
| Outcome | Cases |
|---|---|
| Both pass | 140 |
| Only candidate passes | 20 |
| Only baseline passes | 10 |
| Both fail | 30 |
Candidate success is 160/200 = 80%; baseline is 150/200 = 75%. The paired improvement is (20−10)/200 = 5 percentage points. The disagreements carry the comparison: 30 cases changed outcome, with 20 wins and 10 losses.
Let each paired difference d be +1, −1 or 0. Here the approximate standard error of the mean difference is 0.0272, using the sample variance of d. A simple normal 95% interval is about −0.3 to +10.3 percentage points. It crosses zero, so the observed five-point lead is not precise evidence of improvement at that interval level. This normal approximation can be poor for small or rare-event samples; use an appropriate paired method for the actual metric.
Bootstrap or resample the independent unit. If several turns belong to one conversation, resample conversations, not individual turns. Report important slices, account for repeated comparisons, and choose a practically meaningful change before inspecting results. Statistical significance alone does not establish product value.
Build an evaluation set that resembles the job
Use representative held-out requests and deliberate difficult slices. Version the sample definition, labels/rubrics and reference documents. Keep a development set for iteration and a less frequently used release holdout; repeatedly tuning against a test set makes it development data even without gradient training.
For a support assistant, useful cases include a correct policy answer, an ambiguous request needing clarification, an unavailable answer needing escalation, a multi-turn correction, and a retrieved document containing misleading instructions. Score task resolution, factual support, appropriate escalation and operational cost separately. A polite unsupported answer should not pass because its tone is good.
Choose a scorer you can audit
Use deterministic checks when the outcome permits them: schema validation, exact numerical tolerances, code execution or a verified final state. These checks are only as complete as their specification and test coverage.
An LLM judge can apply rubrics to open-ended responses, but it is another noisy model. Evaluate its agreement and error patterns against blinded human annotations; humans also disagree. For pairwise judging, randomize and swap response order, allow ties/abstention, and inspect position, length and model-family biases. Multiple judges may share the same blind spots. Test whether untrusted answer text can manipulate the judge.
Pin the judge model, prompt, decoding and rubric. Store a concise rationale and evidence where available, not a presumed private reasoning trace. A judge change is an evaluator change and must be measured separately from a candidate-model change.
Locate RAG failures
Retrieval metrics need relevance labels: hit rate@k, recall@k, reciprocal rank or NDCG measure different aspects of the returned set/order. Context precision asks how much retrieved material is relevant under the chosen labels; it is not simply “was this text used in the answer?”
Generation metrics include correctness, supported claims, citation accuracy, completeness and appropriate abstention. Compare generation with retrieved context against generation with verified oracle context to isolate retrieval limits. An answer may faithfully repeat an outdated source and still be wrong. Judge-based tools help automate measurements, but reference-free scoring cannot supply missing ground truth for every criterion.
Use public benchmarks for their actual scope
MMLU samples multiple-choice knowledge across subjects; code tests evaluate behavior under their test suites; repository tasks evaluate a model together with tools, environment and time budget. Preference arenas measure choices by a particular population under a particular protocol. Their rankings are not universal measures of usefulness.
As scores approach a benchmark's ceiling, small differences can be less informative. Do not replace that observation with an undated “frontier score.” Check the benchmark version, prompting, sampling, budget and uncertainty. Public data can leak into training, fine-tuning or repeated model selection. Corpus overlap and behavior probes provide evidence of exposure; no single black-box test proves contamination or its effect on a score.
Turn evaluation into a release process
A harness versions tasks, executes the model, records per-example outputs, applies scorers and computes comparable reports. lm-evaluation-harness supplies task/model adapters and scoring for language-model evaluation. Inspect supports configurable evaluations including tool-using agents and sandboxed tasks. Their capabilities overlap; Inspect is not a successor that replaces lm-eval.
Use a small smoke suite to catch broken runs, then the relevant full evaluation. Separate infrastructure errors from model failures with explicit denominator rules. Candidate/baseline runs should share cases, configurations and evaluation conditions. Shadow execution can reveal production-distribution issues without serving candidate answers; it cannot observe user reactions to those answers. A controlled online experiment adds outcome evidence but needs correct randomization, power, instrumentation and guardrails.
Check yourself
The candidate wins on average but fails more billing requests, a critical product slice. Can the average alone justify release?
Solution: No. Inspect the billing sample size, labels and paired failures, then compare the slice to its predeclared acceptance requirement. A broad improvement can coexist with an unacceptable local regression. More data may be needed; an average cannot decide the requirement for you.
Where to go next
RAG gives a system to evaluate component by component. Harness Engineering turns acceptance criteria and execution evidence into an agent workflow.
References
- HELM: standardized, multi-metric language-model evaluation.
- lm-evaluation-harness: task definitions, model adapters and reproducibility details.
- Inspect documentation: evaluation tasks, solvers, scorers and environments.
- NIST confidence intervals: uncertainty for proportions and assumptions behind interval methods.