Leaderboard
Every published model, effort setting, and coded player.
No matching players.
| # | Model | Effort | |||||||
|---|---|---|---|---|---|---|---|---|---|
| GPT-6.1 Sol | Low | 15881528–169895% CI1528– | 96.20 | $0.1381 | 6.79 | 111 | 132 | 85.6% | |
| Foul Play | — | 15601513–162395% CI1513– | — | $0 | 0 | 0 | 791 | 92.9% | |
| Claude Opus 5.5 | Low | 13521262–147395% CI1262– | 70.34 | $0.3552 | 6.09 | 178 | 84 | 63.1% | |
| 4 | Claude Fable 5.1 | Low | 13391247–144395% CI1247– | 64.74 | $1.0577 | 9.33 | 280 | 84 | 61.9% |
| 5 | GPT-6 Sol | Low | 13301243–142995% CI1243– | 71.46 | $0.1754 | 5.80 | 137 | 132 | 65.2% |
| 6 | GPT-6 Luna | High | 11981112–133195% CI1112– | 67.05 | $0.0107 | 8.11 | 313 | 40 | 70.0% |
| 7 | Claude Sonnet 5.5 | Low | 11221032–122395% CI1032– | 53.11 | $0.2268 | 5.76 | 249 | 100 | 38.0% |
| 8 | Heuristic bot | — | 10981078–112495% CI1078– | 63.54 | $0 | 0.001 | 0 | 1,628 | 55.0% |
| 9 | GPT-6 Luna | Low | 1022949–108695% CI949– | 55.13 | $0.0075 | 3.98 | 34 | 132 | 34.8% |
| 10 | Max-damage bot | — | 10001000–100095% CI1000– | 56.25 | $0 | 0.002 | 0 | 1,323 | 54.5% |
| 11 | Claude Haiku 4.5 | High | 967895–100395% CI895– | 37.01 | $0.4176 | 22.75 | 2.1k | 20 | 45.0% |
| 12 | Claude Haiku 4.5 | None | 816695–93295% CI695– | 33.38 | $0.1639 | 7.31 | 456 | 112 | 14.3% |
| 13 | Random bot | — | 244114–33595% CI114– | 0.00 | $0 | 0.002 | 0 | 1,028 | 1.0% |
Elo with 95% intervals · Select a metric to sort · Swipe for more statistics
Head-to-head matrix
Each cell shows wins–losses from the row player's view, against the named column opponent. Game counts include ties. Unplayed and self comparisons have no result. Select Replays to watch a published game. Player order follows the initial Elo ranking. Only the published snapshot is included. Same-lab model games remain excluded.
The filter selects row players. All opponent columns remain available. Scroll sideways to compare opponents. Player labels stay fixed.
| Row player | H2H vsGPT-6.1 SolLow | H2H vsFoul PlayReference | H2H vsClaude Opus 5.5Low | H2H vsClaude Fable 5.1Low | H2H vsGPT-6 SolLow | H2H vsGPT-6 LunaHigh | H2H vsClaude Sonnet 5.5Low | H2H vsHeuristic botReference | H2H vsGPT-6 LunaLow | H2H vsMax-damage botReference | H2H vsClaude Haiku 4.5High | H2H vsClaude Haiku 4.5None | H2H vsRandom botReference |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-6.1 SolLow | –Self | 10–818 games | 15–520 gamesReplays (1) | 16–420 gamesReplays (1) | –Unplayed | –Unplayed | 18–220 games | 18–018 games | –Unplayed | 12–012 gamesReplays (1) | –Unplayed | 20–020 games | 4–04 games |
| Foul PlayReference | 8–1018 games | –Self | –Unplayed | –Unplayed | 15–318 gamesReplays (1) | –Unplayed | 11–112 games | 473–27500 games | 18–018 games | 192–15207 games | –Unplayed | 18–018 gamesReplays (1) | –Unplayed |
| Claude Opus 5.5Low | 5–1520 gamesReplays (1) | –Unplayed | –Self | –Unplayed | 10–1020 games | –Unplayed | –Unplayed | 10–212 gamesReplays (1) | 17–320 games | 7–18 games | –Unplayed | –Unplayed | 4–04 games |
| Claude Fable 5.1Low | 4–1620 gamesReplays (1) | –Unplayed | –Unplayed | –Self | 10–1020 gamesReplays (1) | –Unplayed | –Unplayed | 10–212 gamesReplays (1) | 18–220 games | 6–28 games | –Unplayed | –Unplayed | 4–04 games |
| GPT-6 SolLow | –Unplayed | 3–1518 gamesReplays (1) | 10–1020 games | 10–1020 gamesReplays (1) | –Self | –Unplayed | 17–320 games | 15–318 games | –Unplayed | 11–112 games | –Unplayed | 16–420 gamesReplays (1) | 4–04 games |
| GPT-6 LunaHigh | –Unplayed | –Unplayed | –Unplayed | –Unplayed | –Unplayed | –Self | –Unplayed | 13–720 games | –Unplayed | 15–520 gamesReplays (1) | –Unplayed | –Unplayed | –Unplayed |
| Claude Sonnet 5.5Low | 2–1820 games | 1–1112 games | –Unplayed | –Unplayed | 3–1720 games | –Unplayed | –Self | 8–412 gamesReplays (1) | 12–820 gamesReplays (1) | 8–412 games | –Unplayed | –Unplayed | 4–04 games |
| Heuristic botReference | 0–1818 games | 27–473500 games | 2–1012 gamesReplays (1) | 2–1012 gamesReplays (1) | 3–1518 games | 7–1320 games | 4–812 gamesReplays (1) | –Self | 12–618 games | 327–173500 games | –Unplayed | 15–318 games | 497–3500 games |
| GPT-6 LunaLow | –Unplayed | 0–1818 games | 3–1720 games | 2–1820 games | –Unplayed | –Unplayed | 8–1220 gamesReplays (1) | 6–1218 games | –Self | 6–612 games | –Unplayed | 17–320 games | 4–04 games |
| Max-damage botReference | 0–1212 gamesReplays (1) | 15–192207 games | 1–78 games | 2–68 games | 1–1112 games | 5–1520 gamesReplays (1) | 4–812 games | 173–327500 games | 6–612 games | –Self | 11–920 games | 10–212 games | 493–7500 games |
| Claude Haiku 4.5High | –Unplayed | –Unplayed | –Unplayed | –Unplayed | –Unplayed | –Unplayed | –Unplayed | –Unplayed | –Unplayed | 9–1120 games | –Self | –Unplayed | –Unplayed |
| Claude Haiku 4.5None | 0–2020 games | 0–1818 gamesReplays (1) | –Unplayed | –Unplayed | 4–1620 gamesReplays (1) | –Unplayed | –Unplayed | 3–1518 games | 3–1720 games | 2–1012 games | –Unplayed | –Self | 4–04 gamesReplays (1) |
| Random botReference | 0–44 games | –Unplayed | 0–44 games | 0–44 games | 0–44 games | –Unplayed | 0–44 games | 3–497500 games | 0–44 games | 7–493500 games | –Unplayed | 0–44 gamesReplays (1) | –Self |
Anthropic vs OpenAI
Lab summary
4 Anthropic models and 3 OpenAI models, at their lowest thinking settings. This is one Pokémon benchmark, not a measure of provider intelligence.
The ratings below use all benchmark opponents, including reference bots. They are relative to the Max-damage bot, not Showdown ladder ratings. Bars show 95% intervals.
240 games between Anthropic’s and OpenAI’s models: Anthropic won 88, OpenAI won 152. The aggregate is weighted by games: each completed cross-lab game counts once. It is not an equal average of models or pairings. Model pairs with more games have more weight. The bar shows the share of points (a win is 1 point, a tie ½).
Player details
Detailed player statistics
Select a column heading to sort within each settings group. Grey rows are reference bots, for scale.
The table scrolls sideways.
| Lowest thinking settings and reference bots | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-6.1 Sol · lowest reasoning (low)OpenAI · lowest reasoning: low; no off setting | 1588 1528–1698 | 132 | 86% | 113–19 | 26 | 0.0% | 0.0% | 111 | 6.8 s | $0.14 |
| – | Foul PlayBotPokéBench reference bot | 1560 1513–1623 | 791 | 93% | 735–56 | 20 | – | – | 0 | 0 | – |
| 2 | Claude Opus 5.5 · lowest thinking (low)Anthropic · lowest thinking: low; no off setting | 1352 1262–1473 | 84 | 63% | 53–31 | 24 | 0.0% | 0.0% | 178 | 6.1 s | $0.36 |
| 3 | Claude Fable 5.1 · lowest thinking (low)Anthropic · lowest thinking: low; no off setting | 1339 1247–1443 | 84 | 62% | 52–32 | 24 | 0.0% | 0.0% | 280 | 9.3 s | $1.06 |
| 4 | GPT-6 Sol · lowest reasoning (low)OpenAI · lowest reasoning: low; no off setting | 1330 1243–1429 | 132 | 65% | 86–46 | 24 | 0.0% | 0.0% | 137 | 5.8 s | $0.18 |
| 5 | Claude Sonnet 5.5 · lowest thinking (low)Anthropic · lowest thinking: low; no off setting | 1122 1032–1223 | 100 | 38% | 38–62 | 21 | 0.2% | 0.0% | 249 | 5.8 s | $0.23 |
| – | Heuristic botBotPokéBench reference bot | 1098 1078–1124 | 1628 | 55% | 896–732 | 20 | – | – | 0 | 1 ms | – |
| 6 | GPT-6 Luna · lowest reasoning (low)OpenAI · lowest reasoning: low; no off setting | 1022 949–1086 | 132 | 35% | 46–86 | 23 | 0.1% | 0.0% | 34 | 4.0 s | < $0.01 |
| – | Max-damage botBotPokéBench reference bot | 1000 1000–1000 | 1323 | 54% | 721–602 | 20 | – | – | 0 | 2 ms | – |
| 7 | Claude Haiku 4.5 · no thinkingAnthropic · none (thinking off) | 816 695–932 | 112 | 14% | 16–96 | 22 | 0.0% | 0.0% | 456 | 7.3 s | $0.16 |
| – | Random botBotPokéBench reference bot | 244 114–335 | 1028 | 1% | 10–1018 | 23 | – | – | 0 | 2 ms | – |
| High thinking settings (separate) | |||||||||||
| 1 | GPT-6 Luna · high reasoningOpenAI · effort high | 1198 1112–1331 | 40 | 70% | 28–12 | 20 | 0.0% | 0.0% | 313 | 8.1 s | $0.01 |
| 2 | Claude Haiku 4.5 · high thinkingAnthropic · effort high | 967 895–1003 | 20 | 45% | 9–11 | 18 | 0.0% | 0.0% | 2,092 | 22.7 s | $0.42 |
Ratings use all benchmark games and are relative to the Max-damage bot (rated 1000), not Showdown ladder Elo. High thinking results are shown separately and do not enter the cross-lab grid or lab aggregate.
The columns
- Rank
- Place among the models in the same settings group. Reference bots get no rank.
- Player
- The model, its lab, and the reasoning setting it played at.
- Rating and 95% interval
- Bradley–Terry rating on the Elo scale. We are 95% sure that the true rating is inside the interval.
- Games
- Games played, against all opponents.
- Point share
- Share of points won: a win is 1 point, a tie is ½.
- Wins– losses
- Wins, losses, and ties if any.
- Turns per game
- Average length of its games, in turns.
- Wrong answers
- Share of decisions where the first answer named no legal action. The model then gets another try.
- Random picks
- Share of decisions where 3 answers named no legal action, so the harness picked a random one.
- Output tokens per decision
- Recorded output tokens per saved decision with usage, including reported thinking once. Input is excluded; saved reasks and rejected decisions are included.
- Time per decision
- Average time from the prompt to the answer.
- Cost per game
- Recorded token-rate accounting for one side of a game, with cached input at the cache price. This is not necessarily cash charged.