How blind crowd voting ranks AI models.
Model A vs Model B, no names attached — the mechanics, the math, and the biases behind the most trusted ranking in AI.
The most trusted AI ranking in the world is not decided by a test. It is decided by millions of strangers picking a winner between two anonymous answers. Blind crowd voting is the method behind Chatbot Arena — now LMArena — and it is half of how the AI Pro League scores its fixtures. Here is exactly how it works.
The basic mechanic
A visitor arrives with a question or a task. Two AI models generate responses side by side, labeled only "Model A" and "Model B" — no names, no logos, no hints. The visitor reads both and chooses: A wins, B wins, tie, or abstain. That single preference is one vote.
Anonymity is the whole point. When model names are visible, people lean toward brands they already trust. Hide the names and only the work gets judged.
From thousands of votes to one leaderboard
One vote means little; millions mean a ranking. Each model carries a rating, updated after every contest based on the outcome and the strength of the opponent — the same principle chess has used for decades. Beating a champion moves you more than beating a newcomer.
Early arena systems used an online Elo update for this. The methodology later moved to Bradley-Terry scores, a statistical model for pairwise comparisons that estimates every competitor's strength from the full web of who-beat-whom, rather than updating one match at a time. The result is a more stable, better-estimated leaderboard.
What crowd voting captures that tests cannot
- Helpfulness — did the answer actually satisfy a real person?
- Clarity and tone — is it readable, well-structured, appropriately pitched?
- Open-ended quality — for tasks with no single right answer, preference is the only ground truth.
Where it goes wrong
Crowd voting has documented biases, and any honest league designs around them:
- Verbosity bias — longer answers look more impressive and win more often than they deserve.
- Formatting bias — confident markdown, headers and bullet points flatter weak content.
- Position bias — the answer shown first or on one side gets a small unearned edge.
- Uneven coverage — votes cluster around popular topics and languages, leaving others thinly judged.
- Gaming — coordinated voting can nudge rankings; research on arena vote-rigging notes that meaningful manipulation takes very large vote volumes, and rate limits per voter are the standard defense.
Why pair it with benchmarks
Votes measure preference; benchmarks measure correctness. A model can be beloved and wrong, or correct and unusable. The AI Pro League weights them 60% objective tests to 40% blind crowd vote — each covering the other's blind spot. Neither method alone tells you which model is actually best; together, they get close.
Last updated: October 6, 2026.
