Clef* None Open weights

CloudflareReasoning setting: none (typed decision)Open weights

Cloudflare's decision model, called through OpenRouter. It plays with a short-context prompt: its route reads only about the first 2,000 tokens, so each decision gets a fresh summary of the battle (no battle log), with the set information of every opponent Pokémon seen so far.

Rating
70395% interval 595⁠–⁠793
Rank
13of 14 players, by Elo
Win rate
18%29–131 in 160 games
Turns / game
22.2
Asked again
0.0%decisions where the first answer had no legal choice
Invalid decisions
0.0%three failed answers, then the first legal action
Tokens / turn
0output and thinking tokens
Price / game
< $0.01retries included
Seconds / turn
445 msper accepted decision

Its rating among all players

Against each opponent

Point share: the share of points won (a win is 1, a tie is ½). Expected: the share that the two ratings predict. A large gap between them can mean that this pairing does not follow what one rating per player predicts.

OpponentRatingGamesPointsPoint shareExpected
Foul Play17281000%0%
GPT-6.1 Sol (low)15602000%1%
GPT-6 Sol (low)14182000%2%
Claude Sonnet 5.5 (low)11342000%8%
Heuristic bot111310110%9%
GPT-6 Luna (low)103820525%13%
Max-damage bot100010220%15%
Claude Haiku 4.5 (no reasoning)84520525%31%
JEV 1.13 (no reasoning)81620735%34%
Random bot27310990%92%

Replays

Back to the leaderboard