Gen 9 Random Battle

PokéBench

AI models and coded players, ranked by benchmark Elo.

Showing 13 of 13
#ModelEffort
Elo rank 1: Master Ball1GPT-6.1 SolOpenAI · LowLow15881528–1698
95% CI1528–1698
96.20$0.13816.7911113285.6%
Elo rank 2: Ultra Ball2Foul PlayOther—15601513–1623
95% CI1513–1623
—$00079192.9%
Elo rank 3: Poké Ball3Claude Opus 5.5Anthropic · LowLow13521262–1473
95% CI1262–1473
70.34$0.35526.091788463.1%
4Claude Fable 5.1Anthropic · LowLow13391247–1443
95% CI1247–1443
64.74$1.05779.332808461.9%
5GPT-6 SolOpenAI · LowLow13301243–1429
95% CI1243–1429
71.46$0.17545.8013713265.2%
6GPT-6 LunaOpenAI · HighHigh11981112–1331
95% CI1112–1331
67.05$0.01078.113134070.0%
7Claude Sonnet 5.5Anthropic · LowLow11221032–1223
95% CI1032–1223
53.11$0.22685.7624910038.0%
8Heuristic botOther—10981078–1124
95% CI1078–1124
63.54$00.00101,62855.0%
9GPT-6 LunaOpenAI · LowLow1022949–1086
95% CI949–1086
55.13$0.00753.983413234.8%
10Max-damage botOther—10001000–1000
95% CI1000–1000
56.25$00.00201,32354.5%
11Claude Haiku 4.5Anthropic · HighHigh967895–1003
95% CI895–1003
37.01$0.417622.752.1k2045.0%
12Claude Haiku 4.5Anthropic · NoneNone816695–932
95% CI695–932
33.38$0.16397.3145611214.3%
13Random botOther—244114–335
95% CI114–335
0.00$00.00201,0281.0%

Elo with 95% intervals · Select a metric to sort · Swipe for more statistics

Head-to-head matrix

Each cell shows wins–losses from the row player's view, against the named column opponent. Game counts include ties. Unplayed and self comparisons have no result. Select Replays to watch a published game. Player order follows the initial Elo ranking. Only the published snapshot is included. Same-lab model games remain excluded.

The filter selects row players. All opponent columns remain available. Scroll sideways to compare opponents. Player labels stay fixed.

Row playerH2H vsGPT-6.1 SolLowH2H vsFoul PlayReferenceH2H vsClaude Opus 5.5LowH2H vsClaude Fable 5.1LowH2H vsGPT-6 SolLowH2H vsGPT-6 LunaHighH2H vsClaude Sonnet 5.5LowH2H vsHeuristic botReferenceH2H vsGPT-6 LunaLowH2H vsMax-damage botReferenceH2H vsClaude Haiku 4.5HighH2H vsClaude Haiku 4.5NoneH2H vsRandom botReference
GPT-6.1 SolLow–Self10–818 games15–520 games
Replays (1)
16–420 games
Replays (1)
–Unplayed–Unplayed18–220 games18–018 games–Unplayed12–012 games
Replays (1)
–Unplayed20–020 games4–04 games
Foul PlayReference8–1018 games–Self–Unplayed–Unplayed15–318 games
Replays (1)
–Unplayed11–112 games473–27500 games
Replays (11)
18–018 games192–15207 games
Replays (4)
–Unplayed18–018 games
Replays (1)
–Unplayed
Claude Opus 5.5Low5–1520 games
Replays (1)
–Unplayed–Self–Unplayed10–1020 games–Unplayed–Unplayed10–212 games
Replays (1)
17–320 games7–18 games–Unplayed–Unplayed4–04 games
Claude Fable 5.1Low4–1620 games
Replays (1)
–Unplayed–Unplayed–Self10–1020 games
Replays (1)
–Unplayed–Unplayed10–212 games
Replays (1)
18–220 games6–28 games–Unplayed–Unplayed4–04 games
GPT-6 SolLow–Unplayed3–1518 games
Replays (1)
10–1020 games10–1020 games
Replays (1)
–Self–Unplayed17–320 games15–318 games–Unplayed11–112 games–Unplayed16–420 games
Replays (1)
4–04 games
GPT-6 LunaHigh–Unplayed–Unplayed–Unplayed–Unplayed–Unplayed–Self–Unplayed13–720 games–Unplayed15–520 games
Replays (1)
–Unplayed–Unplayed–Unplayed
Claude Sonnet 5.5Low2–1820 games1–1112 games–Unplayed–Unplayed3–1720 games–Unplayed–Self8–412 games
Replays (1)
12–820 games
Replays (1)
8–412 games–Unplayed–Unplayed4–04 games
Heuristic botReference0–1818 games27–473500 games
Replays (11)
2–1012 games
Replays (1)
2–1012 games
Replays (1)
3–1518 games7–1320 games4–812 games
Replays (1)
–Self12–618 games327–173500 games
Replays (11)
–Unplayed15–318 games497–3500 games
Replays (10)
GPT-6 LunaLow–Unplayed0–1818 games3–1720 games2–1820 games–Unplayed–Unplayed8–1220 games
Replays (1)
6–1218 games–Self6–612 games–Unplayed17–320 games4–04 games
Max-damage botReference0–1212 games
Replays (1)
15–192207 games
Replays (4)
1–78 games2–68 games1–1112 games5–1520 games
Replays (1)
4–812 games173–327500 games
Replays (11)
6–612 games–Self11–920 games10–212 games493–7500 games
Replays (11)
Claude Haiku 4.5High–Unplayed–Unplayed–Unplayed–Unplayed–Unplayed–Unplayed–Unplayed–Unplayed–Unplayed9–1120 games–Self–Unplayed–Unplayed
Claude Haiku 4.5None0–2020 games0–1818 games
Replays (1)
–Unplayed–Unplayed4–1620 games
Replays (1)
–Unplayed–Unplayed3–1518 games3–1720 games2–1012 games–Unplayed–Self4–04 games
Replays (1)
Random botReference0–44 games–Unplayed0–44 games0–44 games0–44 games–Unplayed0–44 games3–497500 games
Replays (10)
0–44 games7–493500 games
Replays (11)
–Unplayed0–44 games
Replays (1)
–Self

Anthropic vs OpenAI

Lab summary

4 Anthropic models and 3 OpenAI models, at their lowest thinking settings. This is one Pokémon benchmark, not a measure of provider intelligence.

The ratings below use all benchmark opponents, including reference bots. They are relative to the Max-damage bot, not Showdown ladder ratings. Bars show 95% intervals.

Anthropic

Highest benchmark rating

Claude Opus 5.5 · lowest thinking (low)

1352 95% interval 1262⁠–⁠1473

OpenAI

Highest benchmark rating

GPT-6.1 Sol · lowest reasoning (low)

1588 95% interval 1528⁠–⁠1698

240 games between Anthropic’s and OpenAI’s models: Anthropic won 88, OpenAI won 152. The aggregate is weighted by games: each completed cross-lab game counts once. It is not an equal average of models or pairings. Model pairs with more games have more weight. The bar shows the share of points (a win is 1 point, a tie ½).

Gen 9 Random Battle

How well do AI models play Pokémon?

PokéBench ranks AI models by how well they play Pokémon Showdown random battles. Every model gets the same open prompt and the same teams, and every rating comes with its 95% interval.

GPT-6.1 Sol · lowest reasoning (low) chose Keldeo-Resolute.

“Keldeo resists both of Bisharp’s STAB types and threatens it with a quadruple-effective S…”

Watch this battle
Each model gives a reason for each move. Select the text box to watch the battle.

Method

The method in one screen

Fixed seeds

A seed fixes both teams and every random roll of a battle. Anyone can play the same games again.

Mirror matches

Each pair of teams is played twice, with the players swapped. A lucky team then helps each side once.

Ratings with intervals

One Bradley–Terry model, fitted to all games at once, gives the Elo ratings. 200 bootstrap samples give each 95% interval.

One open prompt

Each game is one conversation, in the same plain text for every model: what a player sees on screen, and the legal actions. No tools, no hints.

The cost of each game

We count the tokens and the price of each answer, so every rating comes with its cost per game.

The full method

The prompt, the rating model, the limits, and how to repeat a run.

Read the method

Enter the game

Walk around PokéBench Town. The League holds the leaderboard, the TV in your house shows battles, and the Professor in the lab explains the method.