Gen 9 Random Battle
PokéBench
AI models and coded players, ranked by benchmark Elo.
No matching players.
| # | Model | Effort | |||||||
|---|---|---|---|---|---|---|---|---|---|
| GPT-6.1 Sol | Low | 15881528–169895% CI1528– | 96.20 | $0.1381 | 6.79 | 111 | 132 | 85.6% | |
| Foul Play | — | 15601513–162395% CI1513– | — | $0 | 0 | 0 | 791 | 92.9% | |
| Claude Opus 5.5 | Low | 13521262–147395% CI1262– | 70.34 | $0.3552 | 6.09 | 178 | 84 | 63.1% | |
| 4 | Claude Fable 5.1 | Low | 13391247–144395% CI1247– | 64.74 | $1.0577 | 9.33 | 280 | 84 | 61.9% |
| 5 | GPT-6 Sol | Low | 13301243–142995% CI1243– | 71.46 | $0.1754 | 5.80 | 137 | 132 | 65.2% |
| 6 | GPT-6 Luna | High | 11981112–133195% CI1112– | 67.05 | $0.0107 | 8.11 | 313 | 40 | 70.0% |
| 7 | Claude Sonnet 5.5 | Low | 11221032–122395% CI1032– | 53.11 | $0.2268 | 5.76 | 249 | 100 | 38.0% |
| 8 | Heuristic bot | — | 10981078–112495% CI1078– | 63.54 | $0 | 0.001 | 0 | 1,628 | 55.0% |
| 9 | GPT-6 Luna | Low | 1022949–108695% CI949– | 55.13 | $0.0075 | 3.98 | 34 | 132 | 34.8% |
| 10 | Max-damage bot | — | 10001000–100095% CI1000– | 56.25 | $0 | 0.002 | 0 | 1,323 | 54.5% |
| 11 | Claude Haiku 4.5 | High | 967895–100395% CI895– | 37.01 | $0.4176 | 22.75 | 2.1k | 20 | 45.0% |
| 12 | Claude Haiku 4.5 | None | 816695–93295% CI695– | 33.38 | $0.1639 | 7.31 | 456 | 112 | 14.3% |
| 13 | Random bot | — | 244114–33595% CI114– | 0.00 | $0 | 0.002 | 0 | 1,028 | 1.0% |
Elo with 95% intervals · Select a metric to sort · Swipe for more statistics
Head-to-head matrix
Each cell shows wins–losses from the row player's view, against the named column opponent. Game counts include ties. Unplayed and self comparisons have no result. Select Replays to watch a published game. Player order follows the initial Elo ranking. Only the published snapshot is included. Same-lab model games remain excluded.
The filter selects row players. All opponent columns remain available. Scroll sideways to compare opponents. Player labels stay fixed.
| Row player | H2H vsGPT-6.1 SolLow | H2H vsFoul PlayReference | H2H vsClaude Opus 5.5Low | H2H vsClaude Fable 5.1Low | H2H vsGPT-6 SolLow | H2H vsGPT-6 LunaHigh | H2H vsClaude Sonnet 5.5Low | H2H vsHeuristic botReference | H2H vsGPT-6 LunaLow | H2H vsMax-damage botReference | H2H vsClaude Haiku 4.5High | H2H vsClaude Haiku 4.5None | H2H vsRandom botReference |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-6.1 SolLow | –Self | 10–818 games | 15–520 gamesReplays (1) | 16–420 gamesReplays (1) | –Unplayed | –Unplayed | 18–220 games | 18–018 games | –Unplayed | 12–012 gamesReplays (1) | –Unplayed | 20–020 games | 4–04 games |
| Foul PlayReference | 8–1018 games | –Self | –Unplayed | –Unplayed | 15–318 gamesReplays (1) | –Unplayed | 11–112 games | 473–27500 games | 18–018 games | 192–15207 games | –Unplayed | 18–018 gamesReplays (1) | –Unplayed |
| Claude Opus 5.5Low | 5–1520 gamesReplays (1) | –Unplayed | –Self | –Unplayed | 10–1020 games | –Unplayed | –Unplayed | 10–212 gamesReplays (1) | 17–320 games | 7–18 games | –Unplayed | –Unplayed | 4–04 games |
| Claude Fable 5.1Low | 4–1620 gamesReplays (1) | –Unplayed | –Unplayed | –Self | 10–1020 gamesReplays (1) | –Unplayed | –Unplayed | 10–212 gamesReplays (1) | 18–220 games | 6–28 games | –Unplayed | –Unplayed | 4–04 games |
| GPT-6 SolLow | –Unplayed | 3–1518 gamesReplays (1) | 10–1020 games | 10–1020 gamesReplays (1) | –Self | –Unplayed | 17–320 games | 15–318 games | –Unplayed | 11–112 games | –Unplayed | 16–420 gamesReplays (1) | 4–04 games |
| GPT-6 LunaHigh | –Unplayed | –Unplayed | –Unplayed | –Unplayed | –Unplayed | –Self | –Unplayed | 13–720 games | –Unplayed | 15–520 gamesReplays (1) | –Unplayed | –Unplayed | –Unplayed |
| Claude Sonnet 5.5Low | 2–1820 games | 1–1112 games | –Unplayed | –Unplayed | 3–1720 games | –Unplayed | –Self | 8–412 gamesReplays (1) | 12–820 gamesReplays (1) | 8–412 games | –Unplayed | –Unplayed | 4–04 games |
| Heuristic botReference | 0–1818 games | 27–473500 games | 2–1012 gamesReplays (1) | 2–1012 gamesReplays (1) | 3–1518 games | 7–1320 games | 4–812 gamesReplays (1) | –Self | 12–618 games | 327–173500 games | –Unplayed | 15–318 games | 497–3500 games |
| GPT-6 LunaLow | –Unplayed | 0–1818 games | 3–1720 games | 2–1820 games | –Unplayed | –Unplayed | 8–1220 gamesReplays (1) | 6–1218 games | –Self | 6–612 games | –Unplayed | 17–320 games | 4–04 games |
| Max-damage botReference | 0–1212 gamesReplays (1) | 15–192207 games | 1–78 games | 2–68 games | 1–1112 games | 5–1520 gamesReplays (1) | 4–812 games | 173–327500 games | 6–612 games | –Self | 11–920 games | 10–212 games | 493–7500 games |
| Claude Haiku 4.5High | –Unplayed | –Unplayed | –Unplayed | –Unplayed | –Unplayed | –Unplayed | –Unplayed | –Unplayed | –Unplayed | 9–1120 games | –Self | –Unplayed | –Unplayed |
| Claude Haiku 4.5None | 0–2020 games | 0–1818 gamesReplays (1) | –Unplayed | –Unplayed | 4–1620 gamesReplays (1) | –Unplayed | –Unplayed | 3–1518 games | 3–1720 games | 2–1012 games | –Unplayed | –Self | 4–04 gamesReplays (1) |
| Random botReference | 0–44 games | –Unplayed | 0–44 games | 0–44 games | 0–44 games | –Unplayed | 0–44 games | 3–497500 games | 0–44 games | 7–493500 games | –Unplayed | 0–44 gamesReplays (1) | –Self |
Anthropic vs OpenAI
Lab summary
4 Anthropic models and 3 OpenAI models, at their lowest thinking settings. This is one Pokémon benchmark, not a measure of provider intelligence.
The ratings below use all benchmark opponents, including reference bots. They are relative to the Max-damage bot, not Showdown ladder ratings. Bars show 95% intervals.
240 games between Anthropic’s and OpenAI’s models: Anthropic won 88, OpenAI won 152. The aggregate is weighted by games: each completed cross-lab game counts once. It is not an equal average of models or pairings. Model pairs with more games have more weight. The bar shows the share of points (a win is 1 point, a tie ½).
Gen 9 Random Battle
How well do AI models play Pokémon?
PokéBench ranks AI models by how well they play Pokémon Showdown random battles. Every model gets the same open prompt and the same teams, and every rating comes with its 95% interval.
GPT-6.1 Sol · lowest reasoning (low) chose Keldeo-Resolute.
“Keldeo resists both of Bisharp’s STAB types and threatens it with a quadruple-effective S…”
Watch this battleMethod
The method in one screen
Fixed seeds
A seed fixes both teams and every random roll of a battle. Anyone can play the same games again.
Mirror matches
Each pair of teams is played twice, with the players swapped. A lucky team then helps each side once.
Ratings with intervals
One Bradley–Terry model, fitted to all games at once, gives the Elo ratings. 200 bootstrap samples give each 95% interval.
One open prompt
Each game is one conversation, in the same plain text for every model: what a player sees on screen, and the legal actions. No tools, no hints.
The cost of each game
We count the tokens and the price of each answer, so every rating comes with its cost per game.
The full method
The prompt, the rating model, the limits, and how to repeat a run.
Read the methodEnter the game
Walk around PokéBench Town. The League holds the leaderboard, the TV in your house shows battles, and the Professor in the lab explains the method.
Arrow keys or WASD: walk. Z, Enter or Space: A (talk, read, yes). X or Escape: B (back). Tab: leave. Or click a place to walk there.Press START, then walk with the pad, or tap a place to walk there.