Every model launch arrives with a chart. New model, competitor models, a row of benchmark names — MMLU, GPQA, SWE-bench — and a set of percentages where the new one is highest.
Those numbers are usually accurate, in the narrow sense that the evaluation really was run and really produced that figure. They are also far less informative than they appear, for three distinct reasons that get conflated constantly.
One of them is uncomfortable enough to state up front: in one widely-used benchmark, researchers found that 57% of the questions they analysed in a subject were wrong.
What a Benchmark Actually Is
A benchmark is a fixed set of questions with known answers. You run a model over it, count the correct responses, and report a percentage.
That design carries an assumption which quietly fails in practice: that the model has not seen the questions before, and that the answers in the key are correct.
Both assumptions turn out to be shaky.
Three Failures, Not One

Note the band. Almost none of this is dishonesty. Contamination is mostly accidental, saturation is what success looks like, and errors in a test are an authoring problem. The numbers mislead without anyone lying.
Failure One: The Test Leaked Into the Training Data
Benchmarks are published openly so that results can be reproduced. Models are trained on enormous scrapes of the open web. You can see the collision coming.
Contamination is when test items, or close variants of them, end up in a model's training data — so the model may be recalling an answer rather than working it out. As Chai and colleagues put it, contaminated test samples "artificially inflate reported performance."
The obvious fix is to find the leaked items and delete them. Their paper explains why that does not really work: removing contaminated items "inevitably alters the evaluation set itself and becomes unreliable when contamination is moderate or severe." Change the test to clean it and you can no longer compare against anyone who used the original.
The alternative — suppressing memorised behaviour at evaluation time — has its own cost, since such interventions "often interfere with normal inference and lead to noticeable performance degradation on clean inputs."
There is no clean solution, which is why the field's response has been to build new benchmarks rather than repair old ones. Microsoft's MMLU-CF is one such rebuild, designed to resist both accidental leakage and deliberate gaming. On that contamination-free version, GPT-4o scored 73.4% five-shot and 71.9% zero-shot — figures notably below what the original MMLU leaderboard had come to suggest.
Failure Two: The Benchmark Itself Is Wrong
This is the finding that should change how you read any MMLU number.
A team led by Aryo Pradipta Gema went through MMLU by hand, re-annotating 5,700 questions across all 57 subjects with a formal error-annotation protocol. Their paper is titled Are We Done with MMLU? and it opens by answering itself: "Maybe not."

Across the whole benchmark they estimate 6.49% of MMLU questions contain errors — roughly one in fifteen. In the Virology subset, 57% of the analysed questions contained errors.
Their conclusion is blunt: MMLU "demonstrates numerous ground truth errors that obscure the true capabilities of LLMs," and re-scoring against the corrected set produced "significant discrepancies with the model performance metrics that were originally reported."
Sit with what a wrong answer key does. A model that reasons correctly and gives the right answer is marked wrong. A model that has memorised the flawed key is marked right. On those items the benchmark is not measuring capability — it is measuring agreement with a mistake, which rewards recall over reasoning. That is the opposite of what it was built to do.
Failure Three: Saturation
The third problem is what happens when a benchmark succeeds too well.
When every frontier model scores in the high eighties or nineties on a test, the remaining gaps stop carrying information. The difference between 91.2% and 92.4% is a handful of questions, and on a test with a 6.49% error rate, a gap that size sits comfortably inside the noise.
This is why leaderboard rankings churn between near-identical models, and why a model can be "state of the art" on Monday by a margin nobody could detect in use. A saturated benchmark has stopped measuring the thing it was built for — not because it broke, but because the models outgrew it.
It is also why new benchmarks keep appearing at greater difficulty. That churn is healthy. It does mean year-on-year comparisons are rarely as clean as a chart implies.
Why Vendor Numbers Still Are Not Lies
It is worth being fair about this, because cynicism is as unhelpful as credulity.
A vendor running a published benchmark and reporting the result is doing something reasonable. The scores are reproducible. The problem is not fabrication — it is that a number from a contaminated, partly-erroneous, saturated test cannot support the weight a launch chart puts on it.
The genuine editorial failure is usually downstream: coverage that reports a two-point gap as a decisive lead. We try to hold to that standard in our own Claude vs ChatGPT comparison, where the honest finding is that flagship capabilities have largely converged and the deciding factors are elsewhere.
How to Read a Benchmark Claim

The fourth point is the practical one and it is worth stating plainly: twenty examples of your own real work, run through two or three candidate models, will tell you more about which to use than any published comparison.
It costs an afternoon. It measures your actual task rather than a proxy. And it is the only evaluation in existence that is guaranteed contamination-free, because your examples have never appeared on the internet for a model to train on.
This is the same instinct that makes people over-trust fluent-sounding AI output in the first place — a polished number, like a polished paragraph, reads as more reliable than it is. Our explainer on AI hallucinations covers that pattern in the models themselves.
The Bottom Line
Benchmarks are useful. They are the only way to compare models at scale, and the field would be far worse off without them.
But treat a published score as weak evidence about a narrow skill, not as a measurement of general capability. Contamination inflates results in ways nobody can fully correct. Real benchmarks contain real errors — 6.49% across MMLU, and a majority of the questions analysed in at least one subject. Saturation has compressed the top of the range until differences fall inside the noise.
When a chart shows your preferred model two points ahead, the honest reading is that the two models are indistinguishable on that test. If the decision matters, run your own. For more on how these systems work underneath, see our LLM hub and our guide to choosing between fine-tuning, RAG and prompting.
Frequently Asked Questions
What is benchmark contamination?
It is when a benchmark's test questions, or close variants of them, appear in a model's training data. Because benchmarks are published openly and models are trained on large web scrapes, this happens easily and usually by accident. The effect is that a model may be recalling an answer rather than deriving it, which artificially inflates its reported score without anyone intending to cheat.
Can't researchers just remove the contaminated questions?
They try, but it creates a new problem. As Chai and colleagues note, deleting contaminated items "inevitably alters the evaluation set itself and becomes unreliable when contamination is moderate or severe" — and a modified test can no longer be compared against published results that used the original. The alternative, suppressing memorised behaviour during evaluation, tends to degrade performance on clean inputs too.
Does MMLU really contain wrong answers?
Yes, and this is well documented. Researchers led by Aryo Pradipta Gema re-annotated 5,700 MMLU questions across all 57 subjects and estimated that 6.49% contain errors — wrong ground-truth answers, questions with more than one correct option, and unclear questions. In the Virology subset, 57% of the questions they analysed contained errors. Re-scoring against the corrected set produced significant discrepancies with originally reported figures.
What does benchmark saturation mean?
It means models score so highly that the remaining differences stop carrying information. When frontier models cluster in the high eighties and nineties, a one- or two-point gap is a handful of questions — well within the noise introduced by the benchmark's own error rate. The benchmark has not broken; the models have outgrown it, which is why harder benchmarks keep being introduced.
Should I ignore benchmark scores entirely?
No. They are the only practical way to compare models at scale and they do carry signal, particularly for large gaps on recent, contamination-resistant tests. Treat a score as weak evidence about one narrow skill rather than a measure of general ability, and be sceptical of small margins. A large gap on a fresh benchmark means something; a two-point gap on a saturated one does not.
What's the best way to choose between models?
Test them on your own work. Collect roughly twenty real examples of the task you actually need — your documents, your questions, your edge cases — and run them through two or three candidates. It takes an afternoon, measures the thing you care about rather than a proxy, and is the only evaluation that is structurally immune to contamination, because your examples have never been published for a model to train on.
Sources
- Gema et al. — Are We Done with MMLU? (arXiv:2406.04127)
- Chai, Zhe & Sakuma — When Benchmarks Leak: Inference-Time Decontamination for LLMs (arXiv:2601.19334)
- Zhao et al. — MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark (arXiv:2412.15194)
- MMLU-Redux dataset and error-annotation protocol (GitHub)
- Hendrycks et al. — Measuring Massive Multitask Language Understanding, the original MMLU paper (arXiv:2009.03300)



