Related research

These are the papers that helped us build our own model. It's coming soon.

Pokémon battles

Human-Level Competitive Pokémon via Scalable Offline Reinforcement Learning with Transformers

Jake Grigsby, Yuqi Xie, Justin Sasek, et al. · 2025

Trains Pokémon agents from reconstructed human battle logs and self-play data using offline reinforcement learning with transformers.

PokaiTrainer: Scaling Equilibrium Search to Competitive Pokémon VGC

Max Yu · 2026

Applies equilibrium search over public beliefs to competitive Pokémon doubles battles with simultaneous actions and stochastic outcomes.

The PokeAgent Challenge: Competitive and Long-Context Learning at Scale

Seth Karten, Jake Grigsby, Tersoo Upaa Jr, et al. · 2026

Introduces a Pokémon benchmark with competitive battling and long-horizon role-playing tracks.

Teamwork under extreme uncertainty: AI for Pokemon ranks 33rd in the world

Nicholas R. Sarantinos · 2022

Presents a Pokémon battle agent that manages team balance and uncertainty through specialized search and evaluation methods.

PokéChamp: an Expert-level Minimax Language Agent

Seth Karten, Andy Luu Nguyen, Chi Jin · 2025

Uses language models to propose moves, model opponents, and estimate values inside minimax search for Pokémon battles.

Showdown AI Competition

Scott Lee, Julian Togelius · 2017

Introduces a Pokémon-based AI competition and compares baseline agents for battles with simultaneous decisions and hidden information.

Winning at Pokémon Random Battles Using Reinforcement Learning

Jett Wang · 2024

This thesis studies Pokémon random battles using Monte Carlo tree search guided by an actor-critic trained through self-play.

Automatic Generation of High-Performance RL Environments

Seth Karten, Rahul Dev Appapogu, Chi Jin · 2026

Studies automated translation and verification of faster reinforcement-learning environments, including Pokémon Showdown.

VGC-Bench: Towards Mastering Diverse Team Strategies in Competitive Pokémon

Cameron Angliss, Jiaxun Cui, Jiaheng Hu, et al. · 2025

Introduces a benchmark for learning and evaluating Pokémon doubles agents across diverse team configurations.

Search and imperfect information

Scalable decision-making for games of imperfect information

Samuel Sokota, Eugene Vinitsky, Hengyuan Hu, et al. · 2026

Introduces Ataraxos and combines self-play reinforcement learning with search for games that contain large amounts of hidden information.

A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games

Samuel Sokota, Ryan D'Orazio, J. Zico Kolter, et al. · 2022

Introduces magnetic mirror descent as a method for reinforcement learning and equilibrium computation in two-player zero-sum games.

The Update-Equivalence Framework for Decision-Time Planning

Samuel Sokota, Gabriele Farina, David J. Wu, et al. · 2023

Builds decision-time search methods by reproducing the updates of learning algorithms under imperfect information.

Abstracting Imperfect Information Away from Two-Player Zero-Sum Games

Samuel Sokota, Ryan D’Orazio, Chun Kai Ling, et al. · 2023

Shows how selected regularized equilibria permit perfect-information formulations of imperfect-information zero-sum games.

General search techniques without common knowledge for imperfect-information games, and application to superhuman Fog of War chess

Brian Hu Zhang, Tuomas Sandholm · 2025

Introduces search methods that avoid common-knowledge enumeration and applies them to Fog of War chess.

Reevaluating Policy Gradient Methods for Imperfect-Information Games

Max Rudolph, Nathan Lichtle, Sobhan Mohammadpour, et al. · 2025

Compares policy-gradient methods with specialized learning methods using exploitability in imperfect-information games.

Student of Games: A unified learning algorithm for both perfect and imperfect information games

Martin Schmid, Matej Moravcik, Neil Burch, et al. · 2021

Combines guided search, self-play learning, and game-theoretic reasoning across perfect- and imperfect-information games.

Combining Deep Reinforcement Learning and Search for Imperfect-Information Games

Noam Brown, Anton Bakhtin, Adam Lerer, et al. · 2020

Introduces ReBeL, which combines reinforcement learning and search over beliefs in imperfect-information games.

Thinking Fast and Slow with Deep Learning and Tree Search

Thomas Anthony, Zheng Tian, David Barber · 2017

Introduces Expert Iteration, where tree search improves a policy and a neural network learns from those improvements.

Monte-Carlo Tree Search as Regularized Policy Optimization

Jean-Bastien Grill, Florent Altché, Yunhao Tang, et al. · 2020

Interprets Monte Carlo tree-search heuristics as approximations to regularized policy optimization.

Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning

Julien Perolat, Bart de Vylder, Daniel Hennes, et al. · 2022

Introduces DeepNash, a Stratego agent trained through model-free multiagent reinforcement learning without search.

GAE Falls Short in Imperfect-Information Self-Play Reinforcement Learning

Zhiyuan Fan, Gabriele Farina · 2026

Introduces a centralized action-value estimator that reduces action-sampling variance in imperfect-information self-play.

Superhuman AI for Generals.io Using Self-Play Reinforcement Learning

Matej Straka, Viliam Lisý, Martin Schmid · 2026

Presents a Generals.io agent trained through self-play with a fast simulator, advantage filtering, and averaged policy parameters.

Sequence models for decisions

AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents

Jake Grigsby, Linxi Fan, Yuke Zhu · 2023

Uses sequence models for reinforcement-learning agents that adapt through context and retain long-term information.

AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers

Jake Grigsby, Justin Sasek, Samyak Parajuli, et al. · 2024

Uses classification-based actor and critic objectives to improve transformer agents across unlabeled tasks with different reward scales.

Amortized Planning with Large-Scale Transformers: A Case Study on Chess

Anian Ruoss, Grégoire Delétang, Sourabh Medapati, et al. · 2024

Studies transformer chess policies trained on annotated games that select moves without explicit search.

Decision Transformer: Reinforcement Learning via Sequence Modeling

Lili Chen, Kevin Lu, Aravind Rajeswaran, et al. · 2021

Formulates offline reinforcement learning as sequence modeling conditioned on desired returns, past states, and actions.

Offline Reinforcement Learning as One Big Sequence Modeling Problem

Michael Janner, Qiyang Li, Sergey Levine · 2021

Models complete trajectories with a transformer and uses beam search to plan from offline data.

Online Decision Transformer

Qinqing Zheng, Amy Zhang, Aditya Grover · 2022

Combines offline pretraining and online fine-tuning of Decision Transformers with entropy regularization.

Multi-Game Decision Transformers

Kuang-Huei Lee, Ofir Nachum, Mengjiao Yang, et al. · 2022

Trains one Decision Transformer offline to play multiple Atari games and studies its scaling behavior.

When should we prefer Decision Transformers for Offline Reinforcement Learning?

Prajjwal Bhargava, Rohan Chitnis, Alborz Geramifard, et al. · 2023

Compares Decision Transformers, behavior cloning, and conservative Q-learning across data quality, task horizon, and stochasticity.

You Can't Count on Luck: Why Decision Transformers and RvS Fail in Stochastic Environments

Keiran Paster, Sheila McIlraith, Jimmy Ba · 2022

Explains failures of return conditioning under randomness and proposes conditioning on average returns of trajectory clusters.

Dichotomy of Control: Separating What You Can Control from What You Cannot

Mengjiao Yang, Dale Schuurmans, Pieter Abbeel, et al. · 2022

Separates controllable decisions from environmental randomness when learning policies conditioned on future outcomes.

Robust Adversarial Reinforcement Learning in Stochastic Games via Sequence Modeling

Xiaohang Tang, Zhuowen Cheng, Satyabrat Kumar · 2025

Conditions transformer policies on NashQ values to improve robustness in adversarial stochastic games.

Maia-2: A Unified Model for Human-AI Alignment in Chess

Zhenwei Tang, Difan Jiao, Reid McIlroy-Young, et al. · 2024

Predicts human chess moves across skill levels using a shared model with skill-aware attention.

Scaling laws

Scaling Scaling Laws with Board Games

Andy L. Jones · 2021

Studies how AlphaZero performance on Hex changes with game size, training compute, and search compute.

Scaling Laws for a Multi-Agent Reinforcement Learning Model

Oren Neumann, Claudius Gros · 2022

Measures how AlphaZero playing strength scales with neural-network size and training compute in Connect Four and Pentago.

AlphaZero Neural Scaling and Zipf's Law: a Tale of Board Games and Power Laws

Oren Neumann, Claudius Gros · 2024

Connects AlphaZero scaling and inverse scaling to the frequency distribution of game states.

Scaling laws for single-agent reinforcement learning

Jacob Hilton, Jie Tang, John Schulman · 2023

Studies how reinforcement-learning performance and optimal model size scale with training compute and environment interactions.

Scaling Laws for Imitation Learning in Single-Agent Games

Jens Tuyls, Dhruv Madeka, Kari Torkkola, et al. · 2023

Studies compute, model, and data scaling for imitation-learning agents in Atari and NetHack.

Training Compute-Optimal Large Language Models

Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, et al. · 2022

Studies the allocation of model size and training data under a fixed language-model training budget.

Chinchilla Scaling: A replication attempt

Tamay Besiroglu, Ege Erdil, Matthew Barnett, et al. · 2024

Reexamines the parametric fitting procedure and confidence intervals in Chinchilla scaling estimates.

Scaling Laws for Neural Language Models

Jared Kaplan, Sam McCandlish, Tom Henighan, et al. · 2020

Measures relationships between language-model loss, model size, dataset size, and training compute.

A Hitchhiker's Guide to Scaling Law Estimation

Leshem Choshen, Yang Zhang, Jacob Andreas · 2024

Studies methods for estimating scaling laws from existing models and intermediate training checkpoints.

Evaluation and uncertainty

Deep Reinforcement Learning at the Edge of the Statistical Precipice

Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, et al. · 2021

Studies uncertainty in reinforcement-learning benchmarks and proposes interval estimates and robust aggregate performance measures.

Deep Reinforcement Learning that Matters

Peter Henderson, Riashat Islam, Philip Bachman, et al. · 2017

Studies how experimental choices and random variation affect reproducibility in deep reinforcement learning.

How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments

Cédric Colas, Olivier Sigaud, Pierre-Yves Oudeyer · 2018

Explains how random-seed counts affect statistical power and error rates in reinforcement-learning comparisons.

A Hitchhiker's Guide to Statistical Comparisons of Reinforcement Learning Algorithms

Cédric Colas, Olivier Sigaud, Pierre-Yves Oudeyer · 2019

Compares statistical tests for reinforcement-learning experiments and examines their robustness to violated assumptions.

Empirical Design in Reinforcement Learning

Andrew Patterson, Samuel Neumann, Martha White, et al. · 2023

Provides methods for experimental design, statistical comparison, and control of experimenter bias in reinforcement learning.

AIVAT: A New Variance Reduction Technique for Agent Evaluation in Imperfect Information Games

Neil Burch, Martin Schmid, Matej Moravcik, et al. · 2018

Introduces an unbiased estimator that reduces evaluation variance from chance events and known player strategies.

Re-evaluating Evaluation

David Balduzzi, Karl Tuyls, Julien Perolat, et al. · 2018

Introduces Nash averaging to reduce bias from redundant tasks or agents in evaluation.

Real World Games Look Like Spinning Tops

Wojciech Marian Czarnecki, Gauthier Gidel, Brendan Tracey, et al. · 2020

Studies transitive strength and strategic cycles in games and relates this structure to training with policy populations.

Approximate exploitability: Learning a best response in large games

Finbarr Timbers, Nolan Bard, Edward Lockhart, et al. · 2020

Estimates exploitability in large games by learning a response that targets an agent’s weaknesses.

Training and optimization

What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Marcin Andrychowicz, Anton Raichuk, Piotr Stańczyk, et al. · 2020

Measures how implementation and design choices affect on-policy reinforcement-learning agents in continuous-control tasks.

Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO

Logan Engstrom, Andrew Ilyas, Shibani Santurkar, et al. · 2020

Studies how implementation details influence the behavior and measured performance of PPO and TRPO.

Batch size-invariance for policy optimization

Jacob Hilton, Karl Cobbe, John Schulman · 2021

Separates proximal and behavior policies to make policy optimization less sensitive to batch size.

An Empirical Model of Large-Batch Training

Sam McCandlish, Jared Kaplan, Dario Amodei, et al. · 2018

Uses gradient noise to estimate useful batch sizes and the tradeoff between compute efficiency and training time.

Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Greg Yang, Edward J. Hu, Igor Babuschkin, et al. · 2022

Introduces a parametrization that transfers tuned hyperparameters from smaller neural networks to larger ones.

Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Alexander Hägele, Elie Bakouch, Atli Kosson, et al. · 2024

Studies constant learning rates with cooldowns as a way to reuse training runs when estimating scaling laws.

Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization

Lisha Li, Kevin Jamieson, Giulia DeSalvo, et al. · 2016

Introduces adaptive allocation of training resources to randomly sampled hyperparameter configurations.

A System for Massively Parallel Hyperparameter Tuning

Liam Li, Kevin Jamieson, Afshin Rostamizadeh, et al. · 2018

Introduces asynchronous successive halving for parallel hyperparameter tuning with early stopping.

Understanding Short-Horizon Bias in Stochastic Meta-Optimization

Yuhuai Wu, Mengye Ren, Renjie Liao, et al. · 2018

Studies how short optimization horizons bias hyperparameter selection toward small learning rates.

AutoRL Hyperparameter Landscapes

Aditya Mohan, Carolin Benjamins, Konrad Wienecke, et al. · 2023

Measures how reinforcement-learning hyperparameter landscapes change during training.

Population Based Training of Neural Networks

Max Jaderberg, Valentin Dalibard, Simon Osindero, et al. · 2017

Jointly optimizes neural-network weights and hyperparameter schedules through a population of training runs.

Provably Efficient Online Hyperparameter Optimization with Population-Based Bandits

Jack Parker-Holder, Vu Nguyen, Stephen Roberts · 2020

Uses probabilistic models to guide population-based hyperparameter optimization with fewer parallel agents.

Self-play systems

Emergent Complexity via Multi-Agent Competition

Trapit Bansal, Jakub Pachocki, Szymon Sidor, et al. · 2017

Shows how competitive self-play can produce complex behavior and a natural learning curriculum.

Deep Reinforcement Learning from Self-Play in Imperfect-Information Games

Johannes Heinrich, David Silver · 2016

Introduces Neural Fictitious Self-Play for learning approximate equilibria in imperfect-information games.

A Unified Game-Theoretic Approach to Multiagent Reinforcement Learning

Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, et al. · 2017

Introduces a game-theoretic learning framework that trains responses to mixtures of policies and computes policy-selection strategies.

Adversarial Policies Beat Superhuman Go AIs

Tony T. Wang, Adam Gleave, Tom Tseng, et al. · 2022

Trains adversarial Go policies that exploit weaknesses in strong Go agents.

Dota 2 with Large Scale Deep Reinforcement Learning

OpenAI: Christopher Berner, Greg Brockman, Brooke Chan, et al. · 2019

Describes OpenAI Five and the distributed self-play training system used for competitive Dota 2.

Accelerating Self-Play Learning in Go

David J. Wu · 2019

Introduces improvements that accelerate neural-network-guided self-play learning in Go.

Grandmaster level in StarCraft II using multi-agent reinforcement learning

Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, et al. · 2019

Describes AlphaStar’s use of human games and a league of adapting strategies to learn StarCraft II.

Mastering the game of Go without human knowledge

David Silver, Julian Schrittwieser, Karen Simonyan, et al. · 2017

Presents a Go agent that learns through self-play reinforcement learning without human game data.

Outracing champion Gran Turismo drivers with deep reinforcement learning

Peter R. Wurman, Samuel Barrett, Kenta Kawamoto, et al. · 2022

Trains Gran Turismo racing agents through deep reinforcement learning with mixed scenarios and rewards for racing conduct.

Model growth and reuse

Net2Net: Accelerating Learning via Knowledge Transfer

Tianqi Chen, Ian Goodfellow, Jonathon Shlens · 2015

Introduces function-preserving transformations that transfer knowledge into deeper or wider neural networks.

Staged Training for Transformer Language Models

Sheng Shen, Pete Walsh, Kurt Keutzer, et al. · 2022

Studies staged growth of transformer language models while preserving loss and training dynamics.

bert2BERT: Towards Reusable Pretrained Language Models

Cheng Chen, Yichun Yin, Lifeng Shang, et al. · 2021

Transfers parameters from smaller pretrained language models to initialize larger models and reduce training cost.

Learning to Grow Pretrained Models for Efficient Transformer Training

Peihao Wang, Rameswar Panda, Lucas Torroba Hennigen, et al. · 2023

Learns linear growth operators that initialize larger transformers from smaller pretrained models.

Kickstarting Deep Reinforcement Learning

Simon Schmitt, Jonathan J. Hudson, Augustin Zidek, et al. · 2018

Uses trained teacher policies to accelerate reinforcement-learning students while allowing students to exceed their teachers.

Reincarnating Reinforcement Learning: Reusing Prior Computation to Accelerate Progress

Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, et al. · 2022

Studies how existing agents and prior computation can accelerate the development of new reinforcement-learning agents.

On Warm-Starting Neural Network Training

Jordan T. Ash, Ryan P. Adams · 2019

Studies generalization problems in warm-started neural networks and methods that reduce those problems.

The Primacy Bias in Deep Reinforcement Learning

Evgenii Nikishin, Max Schwarzer, Pierluca D'Oro, et al. · 2022

Studies overfitting to early reinforcement-learning experience and tests partial network resets as a remedy.

Understanding plasticity in neural networks

Clare Lyle, Zeyu Zheng, Evgenii Nikishin, et al. · 2023

Studies why neural networks lose their ability to adapt and tests design choices that preserve plasticity.