AI benchmarks, explained.
MMLU, HumanEval, SWE-bench — what the famous AI tests really measure, what a good score means, and where every one of them falls short.
AI benchmarks are the standardized tests the industry uses to compare models. They are quoted everywhere and understood almost nowhere. This guide covers what the major benchmarks actually measure, what a good score means, and where each one falls short — so you can read any leaderboard like an insider.
| Benchmark | What it measures | Status |
|---|---|---|
| MMLU | Breadth of factual and conceptual knowledge — about 15,900 multiple-choice questions across 57 subjects | Saturated above 90% at the frontier; documented label errors |
| HumanEval | Code generation in isolation — 164 hand-written Python problems, pass@k scoring | Saturated; a thin slice of real engineering |
| SWE-bench | End-to-end software engineering — 2,294 real GitHub issues across a dozen Python repositories | Expensive and noisy; SWE-bench Verified is the trusted subset |
| GSM8K | Multi-step arithmetic reasoning — 8,500 grade-school math word problems | Saturated |
| GPQA | Graduate-level science questions designed to resist web lookup | Diamond subset still genuinely hard |
| HellaSwag | Commonsense reasoning via sentence completion | Saturated |
| Humanity’s Last Exam | Roughly 3,000 expert-written questions across domains | The current frontier challenge |
MMLU — the general knowledge exam
What it is: Massive Multitask Language Understanding — about 15,900 multiple-choice questions across 57 subjects, from abstract algebra to clinical medicine.
What it measures: breadth of factual and conceptual knowledge. A strong score means the model has absorbed a lot about a lot of things.
The catch: frontier models now score above 90%, so it barely separates the best from each other. It also carries documented label errors, and multiple-choice format lets models exploit test-taking patterns rather than reasoning. Its harder successor, MMLU-Pro, uses 10 answer choices and more reasoning-heavy questions.
HumanEval — the coding sprint
What it is: 164 hand-written Python programming problems from OpenAI. The model writes a function; hidden unit tests decide pass or fail. Scores use pass@k: generate k attempts, count the problem solved if any attempt passes.
What it measures: whether a model can produce correct code in isolation — the purest test of code generation.
The catch: also largely saturated at the frontier, and writing one function from a docstring is a thin slice of real engineering. It says nothing about debugging, reading existing code, or working inside a large codebase.
SWE-bench — the real engineering test
What it is: 2,294 real GitHub issues drawn from a dozen Python repositories. The model must read the issue, explore the codebase with tools, and produce a patch that resolves it.
What it measures: end-to-end software engineering: comprehension, tool use, and repair — not just code generation.
The catch: it is expensive to run and noisy to score, which is why the community leans on SWE-bench Verified, a human-validated subset with ambiguous tasks filtered out. Newer variants like SWE-bench Pro add contamination resistance by holding back private test repositories.
The supporting cast
- GSM8K — 8,500 grade-school math word problems; tests multi-step arithmetic reasoning. Saturated.
- GPQA — graduate-level science questions designed to resist web lookup; the "Diamond" subset is still genuinely hard.
- HellaSwag — sentence completion for commonsense reasoning; saturated.
- Humanity's Last Exam — roughly 3,000 expert-written questions across domains, built to stay difficult for years. The current frontier challenge.
Four ways benchmarks mislead
Saturation
When top models all score above 90%, the test stops distinguishing between them. A one-point gap on a saturated benchmark is noise; the same gap on a hard one is signal.
Data contamination
If benchmark questions leak into training data, models memorize answers instead of demonstrating ability. Researchers have shown models acing older coding problems while failing on newly added ones — the signature of recall, not reasoning.
Gaming the test
Labs optimize for the tests everyone quotes. A model can dominate HumanEval's Python tasks yet stumble refactoring real JavaScript — because the benchmark was the target, not the territory.
Missing the real world
Benchmarks rarely test speed, cost, reliability, safety, or how a model fits into an actual workflow. The highest-scoring model can still be the wrong choice for production.
Last updated: October 6, 2026.
