AI Model Leaderboard

Style control ?
Every vote is a pairwise match, and a Bradley–Terry regression over all matches estimates each model's strength on an Elo-like scale: the reference model is pinned at 1500, and a 400-point lead means roughly 10× higher odds of being preferred. The small figure below each score is a 95% confidence range. With Style control on, the regression also holds formatting and length constant, so a model can't climb by looking polished or writing more; off ranks on the raw votes.
21 models · 13,313 total votes
Last updated Sep 9, 2026
RankModelRatingWin RateBattles
1
GPT-5.6 Sol
OpenAI
1735
95% CI 1707 – 1763
54.5%1238
2
Claude Fable 5
Anthropic
1727
95% CI 1702 – 1752
56.1%1694
3
Claude Opus 4.6
Anthropic
1725
95% CI 1695 – 1756
56.8%1029
4
Claude Opus 4.7
Anthropic
1720
95% CI 1690 – 1749
56.8%1120
5
GPT-5.6 Terra
OpenAI
1700
95% CI 1673 – 1728
50.7%1240
6
Claude Opus 4.8
Anthropic
1695
95% CI 1667 – 1722
53.2%1228
7
Claude Sonnet 4.6
Anthropic
1694
95% CI 1665 – 1723
53.8%1083
8
Claude Opus 5
Anthropic
1685
95% CI 1657 – 1712
55.3%1233
9
GPT-5.6 Luna
OpenAI
1684
95% CI 1656 – 1712
48.5%1203
10
GPT-5.2 Chat
OpenAI
1647
95% CI 1618 – 1675
41.1%1246
11
GPT-5.4
OpenAI
1641
95% CI 1613 – 1668
46.2%1220
12
Kimi K2.6
Moonshot AI
1632
95% CI 1605 – 1660
44.5%1192
13
Claude Sonnet 5
Anthropic
1623
95% CI 1595 – 1651
43.2%1167
14
Grok 4.3
xAI
1622
95% CI 1594 – 1650
40.6%1196
15
DeepSeek V4 Pro
DeepSeek
1603
95% CI 1574 – 1631
41.6%1091
16
DeepSeek V4 Flash
DeepSeek
1599
95% CI 1572 – 1627
41.4%1158
17
Claude Haiku 4.5
Anthropic
1591
95% CI 1563 – 1618
38.7%1182
18
Mistral Large 3
Mistral
1584
95% CI 1555 – 1613
44.9%1086
19
GPT-4o
OpenAI
1500
34.8%1214
20
Llama 3.3 70B
Meta
1458
95% CI 1426 – 1491
21.5%1108
21
Command A Plus
Cohere
1450
95% CI 1420 – 1479
30.3%1209