AI Model Leaderboard
Every vote is a pairwise match, and a Bradley–Terry regression over all matches estimates each model's strength on an Elo-like scale: the reference model is pinned at 1500, and a 400-point lead means roughly 10× higher odds of being preferred. The small figure below each score is a 95% confidence range. With Style control on, the regression also holds formatting and length constant, so a model can't climb by looking polished or writing more; off ranks on the raw votes.
The same Bradley–Terry regression, but a "win" is holding your attention longer than another anonymous answer in the same turn — with an answer's position in the list and its length always regressed out. With Style control on, formatting is held constant as well. Ratings sit on an Elo-like scale: the reference model is pinned at 1500, a 400-point lead means roughly 10× higher odds, and the small figure below each score is a 95% confidence range.
Last updated Sep 9, 2026
| Rank | Model | Rating | Win Rate | Battles |
|---|---|---|---|---|
| 1 | GPT-5.6 Sol OpenAI | 1735 95% CI 1707 – 1763 | 54.5% | 1238 |
| 2 | Claude Fable 5 Anthropic | 1727 95% CI 1702 – 1752 | 56.1% | 1694 |
| 3 | Claude Opus 4.6 Anthropic | 1725 95% CI 1695 – 1756 | 56.8% | 1029 |
| 4 | Claude Opus 4.7 Anthropic | 1720 95% CI 1690 – 1749 | 56.8% | 1120 |
| 5 | GPT-5.6 Terra OpenAI | 1700 95% CI 1673 – 1728 | 50.7% | 1240 |
| 6 | Claude Opus 4.8 Anthropic | 1695 95% CI 1667 – 1722 | 53.2% | 1228 |
| 7 | Claude Sonnet 4.6 Anthropic | 1694 95% CI 1665 – 1723 | 53.8% | 1083 |
| 8 | Claude Opus 5 Anthropic | 1685 95% CI 1657 – 1712 | 55.3% | 1233 |
| 9 | GPT-5.6 Luna OpenAI | 1684 95% CI 1656 – 1712 | 48.5% | 1203 |
| 10 | GPT-5.2 Chat OpenAI | 1647 95% CI 1618 – 1675 | 41.1% | 1246 |
| 11 | GPT-5.4 OpenAI | 1641 95% CI 1613 – 1668 | 46.2% | 1220 |
| 12 | Kimi K2.6 Moonshot AI | 1632 95% CI 1605 – 1660 | 44.5% | 1192 |
| 13 | Claude Sonnet 5 Anthropic | 1623 95% CI 1595 – 1651 | 43.2% | 1167 |
| 14 | Grok 4.3 xAI | 1622 95% CI 1594 – 1650 | 40.6% | 1196 |
| 15 | DeepSeek V4 Pro DeepSeek | 1603 95% CI 1574 – 1631 | 41.6% | 1091 |
| 16 | DeepSeek V4 Flash DeepSeek | 1599 95% CI 1572 – 1627 | 41.4% | 1158 |
| 17 | Claude Haiku 4.5 Anthropic | 1591 95% CI 1563 – 1618 | 38.7% | 1182 |
| 18 | Mistral Large 3 Mistral | 1584 95% CI 1555 – 1613 | 44.9% | 1086 |
| 19 | GPT-4o OpenAI | 1500 | 34.8% | 1214 |
| 20 | Llama 3.3 70B Meta | 1458 95% CI 1426 – 1491 | 21.5% | 1108 |
| 21 | Command A Plus Cohere | 1450 95% CI 1420 – 1479 | 30.3% | 1209 |
?