How AI benchmarks work, and why leaderboard scores can mislead

Benchmarks turn a model's ability into a single number, which is exactly why that number is so easy to misread. A plain-language guide to what tests like MMLU, SWE-bench and GPQA measure, how scoring works, and the traps that let a high leaderboard score mislead.

By Himanshu Sakre

Published

A digital tablet showing a web analytics dashboard with graphs and charts
Photo: weCare Media / Pexels

What a benchmark actually is

An AI benchmark is two things bolted together: a fixed set of questions or tasks, and a rule for scoring the answers. The set is frozen so that every model faces the same test, and the resulting number, usually the percentage of answers a model gets right or the percentage of tasks it completes, is what lands on the leaderboard. It is the machine-learning version of a standardized exam. The point is comparability. If two models sit the same 448 questions under the same rules, the higher score is doing something the lower one is not, at least on those 448 questions.

Three names come up again and again. MMLU, from the paper "Measuring Massive Multitask Language Understanding," is a multiple-choice test spanning 57 subjects, from elementary mathematics to US history, computer science and law. GPQA, "A Graduate-Level Google-Proof Q&A Benchmark," is 448 multiple-choice questions in biology, physics and chemistry, written by domain experts. SWE-bench, "Can Language Models Resolve Real-World GitHub Issues," is 2,294 real software problems drawn from actual GitHub issues and pull requests across popular Python projects.

How a model earns a score

Most benchmarks take one of two shapes. Knowledge tests like MMLU and GPQA are multiple choice: the model picks an option, and the grader counts the fraction it gets right. There is nothing subjective about it, which is exactly the appeal. Task benchmarks like SWE-bench are graded by consequences: the model proposes a code patch, a harness applies it to the real repository and runs that project's own tests, and the issue counts as resolved only if the required tests pass. Either way, the output is a single percentage. These task benchmarks are how the industry measures AI agents, and they are unforgiving. When SWE-bench launched, the strongest model of the day resolved just under 2% of the issues.

A well-built benchmark is hard to fake. GPQA is called "Google-proof" because skilled non-experts scored only 34% even after more than thirty minutes with unrestricted web access, while PhD-level experts in the matching field reached about 65%. A wide gap like that is the sign of a test that measures something real. But the same compression that makes a score easy to rank also makes it easy to misread.

“We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired.”

The Leaderboard Illusion, Singh et al., arXiv, 29 Apr 2025
Three medals on podium blocks symbolizing first, second, and third place awards
A benchmark reduces a model to a score on a fixed test. The number is only as good as the test behind it. Photo: DS stories / Pexels

Why the number can mislead

The first trap is contamination. Benchmarks are usually public, so their questions drift into the next model's training data, and a model can then score well by having effectively seen the answers rather than by reasoning through them. Researchers at Scale AI built GSM1k, a fresh set designed to mirror the well-known GSM8k math benchmark in style and difficulty, precisely to test for this. Evaluated on the new set, leading models showed accuracy drops of up to 8%, and several model families showed signs of systematic overfitting. The researchers also found a positive correlation between how readily a model could reproduce a GSM8k question and how far its score fell, which points to partial memorization. Notably, many frontier models held up well, so contamination is a risk to check for, not a blanket verdict.

The second trap is overfitting to the test itself. Once a benchmark becomes the target every team optimizes for, it stops measuring what it originally did, a pattern old enough to have a name: Goodhart's law. Teams can tune prompts, answer formatting and scaffolding to one specific benchmark until the number climbs without the underlying ability moving much.

The third trap is who is holding the stopwatch. Labs routinely publish their own evaluations, run under configurations they chose, against rival models they set up themselves. An independent study of Chatbot Arena, "The Leaderboard Illusion," documented how private testing lets a few large providers try many variants before release and reveal only the best, identifying 27 private variants that one company tested ahead of a single launch, alongside large gaps in how much arena data different players could access. That is a structural reason to treat a vendor's own chart as a claim rather than a measurement.

The fourth trap is cherry-picking. A vendor can choose which benchmarks to show, which competitor configurations to compare against, and which prompting setup to report, and still present a technically accurate result that is not representative. The number is real; the framing does the work.

How to read a benchmark claim skeptically

Ask five questions before you trust a leaderboard line. Who ran the evaluation, an independent group or the vendor selling the model. Is the test set public and old enough that it may already sit in the training data. What configuration produced the number, and is it one you would actually run. Are there error bars, or is a fraction of a point being sold as a decisive lead. And does the benchmark resemble your task at all, because a model that tops a coding leaderboard tells you little about your support inbox.

The strongest signal is a held-out or fresh result: a score on questions the model provably could not have trained on. When nobody can offer one, treat the claim as a hypothesis rather than a finding.

Our take

Benchmarks are instruments, not verdicts. They earned their place because a shared, scored test beats vibes, and the good ones, built from expert-written questions and objective grading, do track real capability. But a single leaderboard number compresses away contamination, tuning, configuration and the question of who did the measuring. Read the score, then read past it. The people building these models already treat a benchmark result as the start of an evaluation rather than the end of one, and anyone reading their charts should do the same.

Frequently asked questions

What is an AI benchmark?

An AI benchmark is a fixed set of questions or tasks combined with a rule for scoring the answers. The set is frozen so that every model faces the same test, and the leaderboard figure is usually the percentage of answers a model gets right or the percentage of tasks it completes. It works like a standardized exam, and the point is comparability between models.

What do MMLU, GPQA and SWE-bench actually measure?

MMLU is a multiple-choice test spanning 57 subjects, from elementary mathematics to US history, computer science and law. GPQA is 448 graduate-level multiple-choice questions in biology, physics and chemistry, written by domain experts. SWE-bench is 2,294 real software problems drawn from actual GitHub issues and pull requests, and an issue counts as resolved only if the project's own tests pass after the model's patch is applied.

Why is GPQA called Google-proof?

GPQA is called Google-proof because skilled non-experts scored only 34% even after more than thirty minutes with unrestricted web access, while PhD-level experts in the matching field reached about 65%. That wide gap is a sign the test measures something real rather than something you can simply look up.

What is benchmark contamination and how do we know it happens?

Contamination is when a public benchmark's questions drift into a model's training data, so the model can score well by having effectively seen the answers rather than by reasoning through them. Scale AI built GSM1k, a fresh set designed to mirror the GSM8k math benchmark, to test for this, and leading models showed accuracy drops of up to 8%, with several model families showing signs of systematic overfitting. Many frontier models held up well, so contamination is a risk to check for, not a blanket verdict.

Why can a high leaderboard score still mislead?

A single number compresses away several problems: contamination, overfitting to the test, the configuration used, and who ran the evaluation. Labs often publish their own evaluations, and one Chatbot Arena study found a provider privately tested 27 variants before a single launch and revealed only the best. Vendors can also cherry-pick which benchmarks and comparisons to show, presenting a technically accurate result that is not representative.

Sources

What each one is, and whose it is.

  1. 1

    SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, arXiv (Princeton, University of Chicago) (October 10, 2023)

    BenchmarkIndependent of the vendor
  2. PaperIndependent of the vendor
  3. 3

    The Leaderboard Illusion, Cohere Labs and academic co-authors (April 29, 2025)

    PaperIndependent of the vendor