Leaderboard

Every published model, effort setting, and coded player.

Showing 13 of 13
#ModelEffort
Elo rank 1: Master Ball1GPT-6.1 SolOpenAI · LowLow15881528–1698
95% CI1528–1698
96.20$0.13816.7911113285.6%
Elo rank 2: Ultra Ball2Foul PlayOther—15601513–1623
95% CI1513–1623
—$00079192.9%
Elo rank 3: Poké Ball3Claude Opus 5.5Anthropic · LowLow13521262–1473
95% CI1262–1473
70.34$0.35526.091788463.1%
4Claude Fable 5.1Anthropic · LowLow13391247–1443
95% CI1247–1443
64.74$1.05779.332808461.9%
5GPT-6 SolOpenAI · LowLow13301243–1429
95% CI1243–1429
71.46$0.17545.8013713265.2%
6GPT-6 LunaOpenAI · HighHigh11981112–1331
95% CI1112–1331
67.05$0.01078.113134070.0%
7Claude Sonnet 5.5Anthropic · LowLow11221032–1223
95% CI1032–1223
53.11$0.22685.7624910038.0%
8Heuristic botOther—10981078–1124
95% CI1078–1124
63.54$00.00101,62855.0%
9GPT-6 LunaOpenAI · LowLow1022949–1086
95% CI949–1086
55.13$0.00753.983413234.8%
10Max-damage botOther—10001000–1000
95% CI1000–1000
56.25$00.00201,32354.5%
11Claude Haiku 4.5Anthropic · HighHigh967895–1003
95% CI895–1003
37.01$0.417622.752.1k2045.0%
12Claude Haiku 4.5Anthropic · NoneNone816695–932
95% CI695–932
33.38$0.16397.3145611214.3%
13Random botOther—244114–335
95% CI114–335
0.00$00.00201,0281.0%

Elo with 95% intervals · Select a metric to sort · Swipe for more statistics

Head-to-head matrix

Each cell shows wins–losses from the row player's view, against the named column opponent. Game counts include ties. Unplayed and self comparisons have no result. Select Replays to watch a published game. Player order follows the initial Elo ranking. Only the published snapshot is included. Same-lab model games remain excluded.

The filter selects row players. All opponent columns remain available. Scroll sideways to compare opponents. Player labels stay fixed.

Row playerH2H vsGPT-6.1 SolLowH2H vsFoul PlayReferenceH2H vsClaude Opus 5.5LowH2H vsClaude Fable 5.1LowH2H vsGPT-6 SolLowH2H vsGPT-6 LunaHighH2H vsClaude Sonnet 5.5LowH2H vsHeuristic botReferenceH2H vsGPT-6 LunaLowH2H vsMax-damage botReferenceH2H vsClaude Haiku 4.5HighH2H vsClaude Haiku 4.5NoneH2H vsRandom botReference
GPT-6.1 SolLow–Self10–818 games15–520 games
Replays (1)
16–420 games
Replays (1)
–Unplayed–Unplayed18–220 games18–018 games–Unplayed12–012 games
Replays (1)
–Unplayed20–020 games4–04 games
Foul PlayReference8–1018 games–Self–Unplayed–Unplayed15–318 games
Replays (1)
–Unplayed11–112 games473–27500 games
Replays (11)
18–018 games192–15207 games
Replays (4)
–Unplayed18–018 games
Replays (1)
–Unplayed
Claude Opus 5.5Low5–1520 games
Replays (1)
–Unplayed–Self–Unplayed10–1020 games–Unplayed–Unplayed10–212 games
Replays (1)
17–320 games7–18 games–Unplayed–Unplayed4–04 games
Claude Fable 5.1Low4–1620 games
Replays (1)
–Unplayed–Unplayed–Self10–1020 games
Replays (1)
–Unplayed–Unplayed10–212 games
Replays (1)
18–220 games6–28 games–Unplayed–Unplayed4–04 games
GPT-6 SolLow–Unplayed3–1518 games
Replays (1)
10–1020 games10–1020 games
Replays (1)
–Self–Unplayed17–320 games15–318 games–Unplayed11–112 games–Unplayed16–420 games
Replays (1)
4–04 games
GPT-6 LunaHigh–Unplayed–Unplayed–Unplayed–Unplayed–Unplayed–Self–Unplayed13–720 games–Unplayed15–520 games
Replays (1)
–Unplayed–Unplayed–Unplayed
Claude Sonnet 5.5Low2–1820 games1–1112 games–Unplayed–Unplayed3–1720 games–Unplayed–Self8–412 games
Replays (1)
12–820 games
Replays (1)
8–412 games–Unplayed–Unplayed4–04 games
Heuristic botReference0–1818 games27–473500 games
Replays (11)
2–1012 games
Replays (1)
2–1012 games
Replays (1)
3–1518 games7–1320 games4–812 games
Replays (1)
–Self12–618 games327–173500 games
Replays (11)
–Unplayed15–318 games497–3500 games
Replays (10)
GPT-6 LunaLow–Unplayed0–1818 games3–1720 games2–1820 games–Unplayed–Unplayed8–1220 games
Replays (1)
6–1218 games–Self6–612 games–Unplayed17–320 games4–04 games
Max-damage botReference0–1212 games
Replays (1)
15–192207 games
Replays (4)
1–78 games2–68 games1–1112 games5–1520 games
Replays (1)
4–812 games173–327500 games
Replays (11)
6–612 games–Self11–920 games10–212 games493–7500 games
Replays (11)
Claude Haiku 4.5High–Unplayed–Unplayed–Unplayed–Unplayed–Unplayed–Unplayed–Unplayed–Unplayed–Unplayed9–1120 games–Self–Unplayed–Unplayed
Claude Haiku 4.5None0–2020 games0–1818 games
Replays (1)
–Unplayed–Unplayed4–1620 games
Replays (1)
–Unplayed–Unplayed3–1518 games3–1720 games2–1012 games–Unplayed–Self4–04 games
Replays (1)
Random botReference0–44 games–Unplayed0–44 games0–44 games0–44 games–Unplayed0–44 games3–497500 games
Replays (10)
0–44 games7–493500 games
Replays (11)
–Unplayed0–44 games
Replays (1)
–Self

Anthropic vs OpenAI

Lab summary

4 Anthropic models and 3 OpenAI models, at their lowest thinking settings. This is one Pokémon benchmark, not a measure of provider intelligence.

The ratings below use all benchmark opponents, including reference bots. They are relative to the Max-damage bot, not Showdown ladder ratings. Bars show 95% intervals.

Anthropic

Highest benchmark rating

Claude Opus 5.5 · lowest thinking (low)

1352 95% interval 1262⁠–⁠1473

OpenAI

Highest benchmark rating

GPT-6.1 Sol · lowest reasoning (low)

1588 95% interval 1528⁠–⁠1698

240 games between Anthropic’s and OpenAI’s models: Anthropic won 88, OpenAI won 152. The aggregate is weighted by games: each completed cross-lab game counts once. It is not an equal average of models or pairings. Model pairs with more games have more weight. The bar shows the share of points (a win is 1 point, a tie ½).

Player details

Detailed player statistics

Select a column heading to sort within each settings group. Grey rows are reference bots, for scale.

The table scrolls sideways.

Lowest thinking settings and reference bots
1GPT-6.1 Sol · lowest reasoning (low)OpenAI · lowest reasoning: low; no off setting1588 1528⁠–⁠169813286%113–19260.0%0.0%1116.8 s$0.14
–Foul PlayBotPokéBench reference bot1560 1513⁠–⁠162379193%735–5620––00–
2Claude Opus 5.5 · lowest thinking (low)Anthropic · lowest thinking: low; no off setting1352 1262⁠–⁠14738463%53–31240.0%0.0%1786.1 s$0.36
3Claude Fable 5.1 · lowest thinking (low)Anthropic · lowest thinking: low; no off setting1339 1247⁠–⁠14438462%52–32240.0%0.0%2809.3 s$1.06
4GPT-6 Sol · lowest reasoning (low)OpenAI · lowest reasoning: low; no off setting1330 1243⁠–⁠142913265%86–46240.0%0.0%1375.8 s$0.18
5Claude Sonnet 5.5 · lowest thinking (low)Anthropic · lowest thinking: low; no off setting1122 1032⁠–⁠122310038%38–62210.2%0.0%2495.8 s$0.23
–Heuristic botBotPokéBench reference bot1098 1078⁠–⁠1124162855%896–73220––01 ms–
6GPT-6 Luna · lowest reasoning (low)OpenAI · lowest reasoning: low; no off setting1022 949⁠–⁠108613235%46–86230.1%0.0%344.0 s< $0.01
–Max-damage botBotPokéBench reference bot1000 1000⁠–⁠1000132354%721–60220––02 ms–
7Claude Haiku 4.5 · no thinkingAnthropic · none (thinking off)816 695⁠–⁠93211214%16–96220.0%0.0%4567.3 s$0.16
–Random botBotPokéBench reference bot244 114⁠–⁠33510281%10–101823––02 ms–
High thinking settings (separate)
1GPT-6 Luna · high reasoningOpenAI · effort high1198 1112⁠–⁠13314070%28–12200.0%0.0%3138.1 s$0.01
2Claude Haiku 4.5 · high thinkingAnthropic · effort high967 895⁠–⁠10032045%9–11180.0%0.0%2,09222.7 s$0.42

Ratings use all benchmark games and are relative to the Max-damage bot (rated 1000), not Showdown ladder Elo. High thinking results are shown separately and do not enter the cross-lab grid or lab aggregate.

The columns

Rank
Place among the models in the same settings group. Reference bots get no rank.
Player
The model, its lab, and the reasoning setting it played at.
Rating and 95% interval
Bradley–Terry rating on the Elo scale. We are 95% sure that the true rating is inside the interval.
Games
Games played, against all opponents.
Point share
Share of points won: a win is 1 point, a tie is ½.
Wins– losses
Wins, losses, and ties if any.
Turns per game
Average length of its games, in turns.
Wrong answers
Share of decisions where the first answer named no legal action. The model then gets another try.
Random picks
Share of decisions where 3 answers named no legal action, so the harness picked a random one.
Output tokens per decision
Recorded output tokens per saved decision with usage, including reported thinking once. Input is excluded; saved reasks and rejected decisions are included.
Time per decision
Average time from the prompt to the answer.
Cost per game
Recorded token-rate accounting for one side of a game, with cached input at the cache price. This is not necessarily cash charged.