Method
How PokéBench works
This page describes the whole method: the task, what a model sees, how the games are played, how the ratings and their intervals are computed, what a game costs, and the limits of all of it. Each term is explained where it first appears.
Prompt version 3 · Pokémon Showdown commit 3661ce40 · data of 6 October 2026
1. The task
Pokémon Showdown is a free online battle simulator: it plays Pokémon battles by the rules of the games. In a random battle, each player gets a team of six Pokémon chosen at random, each with its level, moves, ability and item already set. Nobody builds a team, so the battle tests how a player plays: when to attack, when to switch, and what the opponent will do next.
PokéBench uses Gen 9 Random Battle, singles: the rules of generation 9 (Pokémon Scarlet and Violet), with one active Pokémon on each side. Each turn, both players choose an action at the same time: a move of their active Pokémon, or a switch to another Pokémon of their team. One time in each battle, a player can also Terastallize one Pokémon, which changes its type. A player wins when all six Pokémon of the opponent have fainted.
Random battles suit a benchmark. There are too many possible teams to learn one good team by heart, and each player sees only part of the opponent’s team, so the game asks for reading and planning, not only for calculation.
2. The harness and the prompt
The harness is the program that runs the games. It plays each game in Showdown’s own simulator, the same code that runs pokemonshowdown.com, at one fixed version. It runs on our own computers: no game is played on the public Showdown server.
A decision is one choice of action: one each turn, and one more when a Pokémon faints and the player must send in another. Each game is one conversation for each model. For each decision, the harness adds one message to the conversation and reads one answer. The model’s own earlier answers stay in the conversation. The conversation has three parts:
- The system prompt: the rules in plain words, the task, and the form of the answer. It is the same for every model and every game.
- The first message: the model’s whole team, what it sees of the opponent, the field, and the numbered legal actions (section 3).
- Each later message: only the new battle events since the model’s last decision, the state now in short form, and the legal actions.
Why one conversation: each provider keeps a prompt cache. When a request starts with text that the provider has seen a short time before, the provider reads that part from its cache, at about a tenth of the normal input price. This works only when the earlier messages do not change, and in one conversation they never change. In the first real games, most of the input tokens came from the cache.
The model first thinks in its own reasoning, which is not part of the answer. Then it answers. The prompt asks for two lines only: ACTION: and the number or the name of an action, then REASON: and one short sentence. The harness reads the last ACTION line and the last REASON line of the answer. The replays show the reason. If an answer names no legal action, the model gets a message with the error and answers again. After 3 answers that name no legal action, the harness picks a legal action at random and counts it as a random pick. If the simulator refuses a choice, for example a switch while the active Pokémon is trapped, the model gets Showdown’s error message and chooses again. At turn 300, a game ends as a tie.
Each model plays at its provider’s normal best reasoning setting (for example “high” effort): how much the model may think before it answers. The harness sets it in every request, and each model’s page shows it. The sides are called “Player 1” and “Player 2”: a model does not know who or what it plays against.
The system prompt (version 3), exactly as every model gets it
You are a player in a Pokémon battle on Pokémon Showdown. The format is [Gen 9] Random Battle, singles. THE RULES - Each player has a team of 6 Pokémon. The game chose each Pokémon at random, with its level, moves, ability, item and Tera type. Weaker Pokémon get higher levels. - Each side has one active Pokémon at a time. - Each turn, both players choose an action at the same time. An action is one move of the active Pokémon, or a switch to a Pokémon on the bench. - One time in the battle, you can Terastallize your active Pokémon when it uses a move. Its type then changes to its Tera type until the battle ends. - When a Pokémon faints, its player sends in a Pokémon from the bench. - You win when all Pokémon of the opponent have fainted. - Showdown's standard clauses apply. For example, Sleep Clause: you cannot make a second Pokémon of the opponent fall asleep. - You do not see the team of the opponent. You see only what the opponent showed in this battle. YOUR TASK You get one message per decision. The first message gives the battle so far, your whole team, what you have seen of the opponent, and a numbered list of your legal actions. Each later message gives the new events, the state now and your legal actions. Choose the action that gives you the best chance to win the battle. ANSWER FORMAT Think as much as you need in your private thinking. Then answer with only these two lines, and no other text: ACTION: <the number of one legal action, or its text exactly as listed> REASON: <one short sentence that tells why>
3. What a model sees
A model sees what a human player sees on Showdown’s battle screen, and nothing more. It gets no damage calculator, no type chart, and no list of the sets that random battles use. The first message of a game has:
- Its own team: for each Pokémon its types, level, HP, status, ability, item, Tera type, stats and moves. For the active Pokémon, the type, power, accuracy and a short description of each move.
- What it has seen of the opponent’s team: species, types, HP as a percentage, status, stat changes, and the moves, ability and item it has seen.
- The hazards and screens on each side, and the weather and terrain.
- The legal actions, numbered.
Each later message has the new events of the battle since the model’s last decision, in the words of Showdown’s battle log, then the state now in short form, and the legal actions.
Here are two real messages from one seeded game: the first message, and a later message at turn 6. A token is a piece of text that models read and write, about four characters of English.
The first message (2,626 characters, about 657 tokens)
BATTLE LOG Battle started between Player 1 and Player 2! Go! Carbink! Player 2 sent out Calyrex (Calyrex-Shadow)! [The opposing Calyrex's As One] The opposing Calyrex has two Abilities! [The opposing Calyrex's Unnerve] Your team is too nervous to eat Berries! == Turn 1 == CURRENT STATE (turn 1) YOUR TEAM (Player 1): - Carbink [ACTIVE], Rock/Fairy, L90, HP 236/236 (100%) Ability: Sturdy, Item: Chesto Berry, Tera type: Fighting Stats: Atk 95, Def 321, SpA 141, SpD 321, Spe 141 Moves: Body Press (Fighting, Physical, 80 BP, 100% accuracy: Uses user's Def stat as Atk in damage calculation.); Iron Defense (Steel, Status, never misses: Raises the user's Defense by 2.); Moonblast (Fairy, Special, 95 BP, 100% accuracy: 30% chance to lower the target's Sp. Atk by 1.); Rest (Psychic, Status, never misses: User sleeps 2 turns and restores HP and status.) - Palkia-Origin [bench], Water/Dragon, L72, HP 249/249 (100%) Ability: Pressure, Item: Lustrous Globe, Tera type: Fire Stats: Atk 149, Def 186, SpA 258, SpD 215, Spe 215 Moves: Spacial Rend; Fire Blast; Hydro Pump; Thunder Wave - Granbull [bench], Fairy, L88, HP 302/302 (100%) Ability: Intimidate, Item: Leftovers, Tera type: Ground Stats: Atk 261, Def 182, SpA 156, SpD 156, Spe 129 Moves: Encore; Play Rough; Earthquake; Thunder Wave - Vivillon-Ocean [bench], Bug/Flying, L83, HP 268/268 (100%) Ability: Compound Eyes, Item: Heavy-Duty Boots, Tera type: Flying Stats: Atk 91, Def 131, SpA 197, SpD 131, Spe 195 Moves: Hurricane; Bug Buzz; Sleep Powder; Quiver Dance - Virizion [bench], Grass/Fighting, L82, HP 283/283 (100%) Ability: Justified, Item: Life Orb, Tera type: Rock Stats: Atk 195, Def 165, SpA 195, SpD 259, Spe 224 Moves: Swords Dance; Leaf Blade; Stone Edge; Close Combat - Weavile [bench], Dark/Ice, L79, HP 239/239 (100%) Ability: Pickpocket, Item: Choice Band, Tera type: Fighting Stats: Atk 235, Def 148, SpA 117, SpD 180, Spe 243 Moves: Ice Shard; Low Kick; Triple Axel; Knock Off OPPONENT'S TEAM (Player 2), as far as you have seen it: - Calyrex-Shadow [ACTIVE], Psychic/Ghost, L64, HP 100% Ability: Unnerve Moves seen: none yet - 5 Pokémon not seen yet YOUR LEGAL ACTIONS 1. move Body Press 2. move Iron Defense 3. move Moonblast 4. move Rest 5. move Body Press + Terastallize (Fighting) 6. move Iron Defense + Terastallize (Fighting) 7. move Moonblast + Terastallize (Fighting) 8. move Rest + Terastallize (Fighting) 9. switch Palkia-Origin 10. switch Granbull 11. switch Vivillon-Ocean 12. switch Virizion 13. switch Weavile Choose one action. Answer with only the ACTION line and the REASON line.
A later message, at turn 6 (1,175 characters, about 294 tokens):
NEW EVENTS The opposing Rotom used Overheat! (Granbull lost 30.8% of its health!) The opposing Rotom's Sp. Atk fell harshly! Granbull used Earthquake! [The opposing Rotom's Levitate] It doesn't affect the opposing Rotom... Granbull restored a little HP using its Leftovers! == Turn 6 == STATE NOW (turn 6) YOUR ACTIVE POKÉMON: - Granbull [ACTIVE], Fairy, L88, HP 227/302 (75%) Moves: Encore; Play Rough; Earthquake; Thunder Wave YOUR BENCH: Palkia-Origin 249/249 (100%); Virizion fainted; Vivillon-Ocean 268/268 (100%); Carbink fainted; Weavile 239/239 (100%) OPPONENT (Player 2): - Rotom-Heat [ACTIVE], Electric/Fire, L83, HP 43% Ability: Levitate Stat changes: spa -4, atk -1 Moves seen: Overheat ALSO SEEN: Calyrex-Shadow fainted; 4 Pokémon not seen yet YOUR LEGAL ACTIONS 1. move Encore 2. move Play Rough 3. move Earthquake 4. move Thunder Wave 5. move Encore + Terastallize (Ground) 6. move Play Rough + Terastallize (Ground) 7. move Earthquake + Terastallize (Ground) 8. move Thunder Wave + Terastallize (Ground) 9. switch Palkia-Origin 10. switch Vivillon-Ocean 11. switch Weavile Choose one action. Answer with only the ACTION line and the REASON line.
4. Seeds and mirror matches
A seed is a short text from which a program draws all of its random numbers: the same seed always gives the same numbers. In PokéBench, seeds fix both teams and every random roll of a battle, such as damage rolls, critical hits and accuracy.
Each run has one seed text. From it, the harness makes the seeds of each pair of games with SHA-256, a hash function: the same input always gives the same output, and different inputs give unrelated outputs. Pair number k gets the same teams in every matchup of the run, so all players face the same luck.
A mirror match is a pair of games with the same seeds and the players swapped. In the first game, player A has team 1 and player B has team 2. In the second game, B has team 1 and A has team 2. A strong team then helps each player once, so most of the luck of the teams cancels out. The rolls during the battle still differ once the players make different choices.
With the same seeds, the same choices always give the same game. A model does not always give the same answer to the same prompt, so its games cannot be played again from the seeds alone. But each game record keeps Showdown’s input log, which replays the game exactly.
5. Ratings and intervals
An Elo rating is a number for the strength of a player. The difference between two ratings gives the chance that one player beats the other. PokéBench uses the Bradley–Terry model, which gives that chance as:
P(A beats B) = 1 / (1 + 10(RB − RA) / 400)
So a player rated 200 points higher is expected to win about 76% of games, and a player rated 400 points higher about 91%, odds of 10 to 1. A tie counts as half a win for each side.
The ratings are fitted to all games at once: they are the ratings that make the results most likely. So, unlike the Elo ratings of a ladder, they do not depend on the order of the games. A weak prior, an assumption made before any game, keeps a rating finite when a player wins or loses every game: each rating starts with a spread (one standard deviation) of 800 points around the average.
The 95% interval shows how sure we are of a rating. We compute it with the bootstrap: we draw a new set of games at random from the games we have, the same number of them, with repeats allowed, and fit the ratings again. We do this 200 times. The middle 95% of a player’s ratings over these samples is its interval. We draw mirror pairs, not single games, because the two games of a pair share their luck. Each sample is shifted by its own Max-damage rating to fix Max-damage at 1000. Its interval is therefore 1000–1000.
More games make an interval shorter. Near a 50% win rate, an interval of about ±50 points needs about 200 games.
6. Reference bots and the rating scale
The models also play reference bots, for scale:
- Foul Play: Foul Play (pmariglia/foul-play, GPL-3.0), a tree-search bot that plays near the top of Showdown's random battle ladder. Here it searches 200 ms per move, much less than on the ladder. It plays as its own program on our local Showdown server.
- Heuristic bot: A faithful port of poke-env 0.16.1's SimpleHeuristicsPlayer, quirks included: it sets entry hazards, switches out of bad matchups, Terastallizes when the type change helps, and otherwise uses its strongest move by a rough damage estimate.
- Max-damage bot: Uses the move that does the most damage by a rough estimate. Switches only when it must.
- Random bot: Chooses one of its legal actions at random.
The heuristic bot is a port of the SimpleHeuristicsPlayer of poke-env (MIT licence). Foul Play by pmariglia (GPL-3.0), at commit 6c467c08, with 200 ms of search per move, plays on our own local Showdown server. Its code is not part of the PokéBench code.
Elo ratings have no natural zero: only the differences between ratings mean something. Our ratings are relative to Max-damage: the scale is set so that the Max-damage bot is rated 1000. These Elo ratings are not directly comparable to Pokémon Showdown ladder ratings. We played no human opponents and set the Max-damage bot’s Elo to 1000. The differences do not depend on this choice: 400 points always mean odds of 10 to 1.
We do not estimate ratings on Pokémon Showdown’s public ladder. No published rating fits our bots at our settings: Foul Play’s 2341 is a peak rating at about 7 seconds per move, and the 1433 of a SimpleHeuristicsPlayer comes from 60 games, read from a figure (Metamon, arXiv 2504.04395).
7. Cost accounting
All model calls go through one OpenAI-compatible API. For each answer, the harness records the tokens: input, cached input, output and reasoning tokens. Cached input is the part of the input that the provider reads from its prompt cache (section 2). The harness prices the tokens at each provider’s list prices, and the cached input at the provider’s cache price. These recorded accounting values are not necessarily cash charges, including calls through subscription-backed gateways.
- Cost per game is the cost of one side of one game.
- Tokens/turn is recorded output tokens per saved decision with usage. Input and cached input are excluded. Reported thinking is included once in output. Saved reasks and simulator-rejected decisions are included. Unreported thinking and unsaved failed-attempt usage are not estimated.
- The raw snapshot retains its original input-plus-output average. The displayed output-only metric comes from a supplement bound to that unchanged snapshot. Cost still includes full input and output usage.
- Time per decision is the time from the request to the answer. It depends on the provider’s load as well as on the model.
8. Threats to validity
- One prompt for all. One plain-text format may suit some models better than others. One format keeps the comparison fair, but a model may play better with another harness.
- Reasoning settings differ. “High” effort does not mean the same amount of thinking at each provider. Each model’s page shows its setting.
- Luck. Mirror matches cancel most of the luck of the teams, but not the luck of the rolls. The intervals include the luck that remains.
- One number per player. The Bradley–Terry model assumes that one number per player describes every pairing. A style can beat another style more often than the ratings say. The head-to-head matrix shows such cases.
- Knowledge from training. A model may have read about random battles, their sets and their strategy. People can read the same pages, so we count this as part of the skill. But it can make a model look stronger than its reasoning alone.
- No time limit. A model can think as long as its setting allows. People on the ladder play against a clock.
- Long conversations. A long game makes a long conversation. A model can play worse late in a long conversation than it would with a short prompt.
- Model versions. A provider can change a model behind the same name. Each run records the model id and the time of each game.
- Wrong answers. Wrong answers and random picks count against a model. Games that end in an error of the harness are left out of the ratings.
9. How to reproduce a run
The harness needs Node 22. From the PokéBench repository:
# Showdown at the pinned commit scripts/setup-showdown.sh cd harness && npm ci # Play the run. It is safe to start it again. npx tsx src/run.ts configs/<run>.json # Ratings and replays, into data/site/ npx tsx src/publish.ts <run> # The prompt that a model sees at turn 6 npx tsx src/prompt-sample.ts 6
A run config sets the seed text, the format, the turn limit and the matchups. With the same config, the bots play exactly the same games. A model’s games can differ, for the reason in section 4. The data on this site comes from these runs: fp-haiku-none, fp-luna-low, fp-sol6-low, fp-sol61-low, fp-sonnet-low, p1-fable-low, p1-haiku-none, p1-luna-low, p1-opus-low, p1-sol6-low, p1-sol61-low, p1-sonnet-low, real-haiku, real-luna, x-fable, x-haiku, x-opus, x-sonnet, x2a-fable, x2a-haiku, x2a-opus, x2a-sonnet, bots-calibration, bots-foulplay.
10. How to cite
@misc{pokebench2026,
title = {Pok{\'e}Bench: How Well Do AI Models Play Pok{\'e}mon?},
author = {{PokéBench}},
year = {2026},
url = {https://pokebench.xyz},
note = {Data of 6 October 2026, prompt version 3}
}