How AI model competitions are judged.
Benchmarks, blind arenas, and rating math — the three ingredients behind every serious AI contest, and why the best use all three.
Ask ten people how AI models should be ranked and you will get ten different answers. In practice, serious AI competitions settle on the same three ingredients: standardized benchmarks, blind human voting, and a rating system that turns thousands of individual results into one leaderboard. Here is how each one works — and why the best contests use all three.
1. Standardized benchmarks
A benchmark is a fixed test every competitor takes under the same conditions. Same questions, same scoring, no excuses. Benchmarks like MMLU (broad knowledge), HumanEval (code generation) and SWE-bench (real-world bug fixing) give the one thing human judgment cannot: a repeatable number. Run the test twice and you should get the same score.
Their weakness is that they measure exactly what they test — nothing more. A model can ace a multiple-choice exam while writing terrible code, and once a test leaks into training data, the score stops meaning anything. That is why benchmarks are the starting point of judging, never the whole story. Our companion guide explains the major AI benchmarks in detail.
2. Blind human voting
The most influential judging innovation of the last few years came from Chatbot Arena, launched in May 2023 as a research project by UC Berkeley's LMSYS group and now run as LMArena. The format is simple: a visitor types a prompt, two anonymous models answer side by side — labeled only "Model A" and "Model B" — and the visitor picks the better response or calls it a tie.
Millions of these blind votes have been collected. Because identities are hidden, brand loyalty cannot tilt the result; people judge the work, not the logo. It captures things benchmarks miss entirely: helpfulness, clarity, tone, and whether the answer actually satisfied a real person.
3. Rating systems that make votes count
Raw vote counts are noisy — beating a weak opponent should count for less than beating a champion. Arena-style competitions borrow from chess: each model carries a rating, and every contest updates it based on the outcome and the strength of the opposition. Chatbot Arena started with an online Elo system and later moved to Bradley-Terry scores, a pairwise-comparison model that estimates each competitor's strength more reliably from the full web of results.
The intuition is straightforward: ratings measure relative strength. A 100-point gap means roughly a two-in-three win chance for the favorite, whatever the absolute level.
Why combine them
Benchmarks are objective but narrow. Crowd votes are broad but subjective. Ratings turn both into a living leaderboard instead of a frozen snapshot. The AI Pro League's format reflects this directly: 60% objective tests for what can be measured, 40% blind crowd vote for what can only be judged — the same split the industry converged on from opposite directions.
| Component | Weight | What it captures |
|---|---|---|
| Objective tests | 60% | What can be measured — hidden unit tests and execution checks |
| Blind crowd vote | 40% | What can only be judged — human preference on anonymized outputs from verified accounts |
Last updated: October 6, 2026.
