12.7 Evaluation science
You can tell a real benchmark result from a marketing one.
Before:08. Large Language Models
Evaluation science is the defence against benchmark theatre: contamination, saturation, and marketing numbers versus held-out truth, with statistical significance separating real gains from sampling noise. It sits late in the frontier module because frontier announcements arrive daily and need filtering. The working conclusion is an internal benchmark you trust — public test data leaks into training sets, so the only numbers beyond suspicion are the ones nobody could have trained on.
Work through these
Benchmark contamination and saturation
Test material leaking into training makes a score meaningless, and popular benchmarks eventually stop discriminating between models. Both problems affect nearly every published comparison.
Held-out and dynamic benchmarks
Benchmarks kept private, and ones that change over time, address the problems above at the cost of being harder to compare against. Understanding the trade helps you read results.
Statistical significance in model comparison
Two models differing by a small margin on a small test set may not differ at all. Applying ordinary statistical reasoning to model comparison is rarer than it should be.
Building an internal benchmark you trust
A benchmark built from your own tasks, kept private, is the only one whose result reliably predicts your outcome. Building it is the practical conclusion of this topic.
Sign in to keep your progress.
Free resources
We haven't checked most of these for screen reader use yet.
Links last checked 29 Aug 2026.
Stuck here?
Ask a mentor. A real person answers, and they can see exactly which topic you're on. Usually within a couple of working days.
Checking your session…
Topics shown in module order.