Related research
These are the papers that helped us build our own model. It's coming soon.
Pokémon battles
Human-Level Competitive Pokémon via Scalable Offline Reinforcement Learning with Transformers
Jake Grigsby, Yuqi Xie, Justin Sasek, et al. · 2025
Trains Pokémon agents from reconstructed human battle logs and self-play data using offline reinforcement learning with transformers.
PokaiTrainer: Scaling Equilibrium Search to Competitive Pokémon VGC
Max Yu · 2026
Applies equilibrium search over public beliefs to competitive Pokémon doubles battles with simultaneous actions and stochastic outcomes.
The PokeAgent Challenge: Competitive and Long-Context Learning at Scale
Seth Karten, Jake Grigsby, Tersoo Upaa Jr, et al. · 2026
Introduces a Pokémon benchmark with competitive battling and long-horizon role-playing tracks.
Teamwork under extreme uncertainty: AI for Pokemon ranks 33rd in the world
Nicholas R. Sarantinos · 2022
Presents a Pokémon battle agent that manages team balance and uncertainty through specialized search and evaluation methods.
PokéChamp: an Expert-level Minimax Language Agent
Seth Karten, Andy Luu Nguyen, Chi Jin · 2025
Uses language models to propose moves, model opponents, and estimate values inside minimax search for Pokémon battles.
Showdown AI Competition
Scott Lee, Julian Togelius · 2017
Introduces a Pokémon-based AI competition and compares baseline agents for battles with simultaneous decisions and hidden information.
Winning at Pokémon Random Battles Using Reinforcement Learning
Jett Wang · 2024
This thesis studies Pokémon random battles using Monte Carlo tree search guided by an actor-critic trained through self-play.
Automatic Generation of High-Performance RL Environments
Seth Karten, Rahul Dev Appapogu, Chi Jin · 2026
Studies automated translation and verification of faster reinforcement-learning environments, including Pokémon Showdown.
VGC-Bench: Towards Mastering Diverse Team Strategies in Competitive Pokémon
Cameron Angliss, Jiaxun Cui, Jiaheng Hu, et al. · 2025
Introduces a benchmark for learning and evaluating Pokémon doubles agents across diverse team configurations.
Search and imperfect information
Scalable decision-making for games of imperfect information
Samuel Sokota, Eugene Vinitsky, Hengyuan Hu, et al. · 2026
Introduces Ataraxos and combines self-play reinforcement learning with search for games that contain large amounts of hidden information.
A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games
Samuel Sokota, Ryan D'Orazio, J. Zico Kolter, et al. · 2022
Introduces magnetic mirror descent as a method for reinforcement learning and equilibrium computation in two-player zero-sum games.
The Update-Equivalence Framework for Decision-Time Planning
Samuel Sokota, Gabriele Farina, David J. Wu, et al. · 2023
Builds decision-time search methods by reproducing the updates of learning algorithms under imperfect information.
Abstracting Imperfect Information Away from Two-Player Zero-Sum Games
Samuel Sokota, Ryan D’Orazio, Chun Kai Ling, et al. · 2023
Shows how selected regularized equilibria permit perfect-information formulations of imperfect-information zero-sum games.
General search techniques without common knowledge for imperfect-information games, and application to superhuman Fog of War chess
Brian Hu Zhang, Tuomas Sandholm · 2025
Introduces search methods that avoid common-knowledge enumeration and applies them to Fog of War chess.
Reevaluating Policy Gradient Methods for Imperfect-Information Games
Max Rudolph, Nathan Lichtle, Sobhan Mohammadpour, et al. · 2025
Compares policy-gradient methods with specialized learning methods using exploitability in imperfect-information games.
Student of Games: A unified learning algorithm for both perfect and imperfect information games
Martin Schmid, Matej Moravcik, Neil Burch, et al. · 2021
Combines guided search, self-play learning, and game-theoretic reasoning across perfect- and imperfect-information games.
Combining Deep Reinforcement Learning and Search for Imperfect-Information Games
Noam Brown, Anton Bakhtin, Adam Lerer, et al. · 2020
Introduces ReBeL, which combines reinforcement learning and search over beliefs in imperfect-information games.
Thinking Fast and Slow with Deep Learning and Tree Search
Thomas Anthony, Zheng Tian, David Barber · 2017
Introduces Expert Iteration, where tree search improves a policy and a neural network learns from those improvements.
Monte-Carlo Tree Search as Regularized Policy Optimization
Jean-Bastien Grill, Florent Altché, Yunhao Tang, et al. · 2020
Interprets Monte Carlo tree-search heuristics as approximations to regularized policy optimization.
Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning
Julien Perolat, Bart de Vylder, Daniel Hennes, et al. · 2022
Introduces DeepNash, a Stratego agent trained through model-free multiagent reinforcement learning without search.
GAE Falls Short in Imperfect-Information Self-Play Reinforcement Learning
Zhiyuan Fan, Gabriele Farina · 2026
Introduces a centralized action-value estimator that reduces action-sampling variance in imperfect-information self-play.
Superhuman AI for Generals.io Using Self-Play Reinforcement Learning
Matej Straka, Viliam Lisý, Martin Schmid · 2026
Presents a Generals.io agent trained through self-play with a fast simulator, advantage filtering, and averaged policy parameters.
Sequence models for decisions
AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents
Jake Grigsby, Linxi Fan, Yuke Zhu · 2023
Uses sequence models for reinforcement-learning agents that adapt through context and retain long-term information.
AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers
Jake Grigsby, Justin Sasek, Samyak Parajuli, et al. · 2024
Uses classification-based actor and critic objectives to improve transformer agents across unlabeled tasks with different reward scales.
Amortized Planning with Large-Scale Transformers: A Case Study on Chess
Anian Ruoss, Grégoire Delétang, Sourabh Medapati, et al. · 2024
Studies transformer chess policies trained on annotated games that select moves without explicit search.
Decision Transformer: Reinforcement Learning via Sequence Modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, et al. · 2021
Formulates offline reinforcement learning as sequence modeling conditioned on desired returns, past states, and actions.
Offline Reinforcement Learning as One Big Sequence Modeling Problem
Michael Janner, Qiyang Li, Sergey Levine · 2021
Models complete trajectories with a transformer and uses beam search to plan from offline data.
Online Decision Transformer
Qinqing Zheng, Amy Zhang, Aditya Grover · 2022
Combines offline pretraining and online fine-tuning of Decision Transformers with entropy regularization.
Multi-Game Decision Transformers
Kuang-Huei Lee, Ofir Nachum, Mengjiao Yang, et al. · 2022
Trains one Decision Transformer offline to play multiple Atari games and studies its scaling behavior.
When should we prefer Decision Transformers for Offline Reinforcement Learning?
Prajjwal Bhargava, Rohan Chitnis, Alborz Geramifard, et al. · 2023
Compares Decision Transformers, behavior cloning, and conservative Q-learning across data quality, task horizon, and stochasticity.
You Can't Count on Luck: Why Decision Transformers and RvS Fail in Stochastic Environments
Keiran Paster, Sheila McIlraith, Jimmy Ba · 2022
Explains failures of return conditioning under randomness and proposes conditioning on average returns of trajectory clusters.
Dichotomy of Control: Separating What You Can Control from What You Cannot
Mengjiao Yang, Dale Schuurmans, Pieter Abbeel, et al. · 2022
Separates controllable decisions from environmental randomness when learning policies conditioned on future outcomes.
Robust Adversarial Reinforcement Learning in Stochastic Games via Sequence Modeling
Xiaohang Tang, Zhuowen Cheng, Satyabrat Kumar · 2025
Conditions transformer policies on NashQ values to improve robustness in adversarial stochastic games.
Maia-2: A Unified Model for Human-AI Alignment in Chess
Zhenwei Tang, Difan Jiao, Reid McIlroy-Young, et al. · 2024
Predicts human chess moves across skill levels using a shared model with skill-aware attention.
Scaling laws
Scaling Scaling Laws with Board Games
Andy L. Jones · 2021
Studies how AlphaZero performance on Hex changes with game size, training compute, and search compute.
Scaling Laws for a Multi-Agent Reinforcement Learning Model
Oren Neumann, Claudius Gros · 2022
Measures how AlphaZero playing strength scales with neural-network size and training compute in Connect Four and Pentago.
AlphaZero Neural Scaling and Zipf's Law: a Tale of Board Games and Power Laws
Oren Neumann, Claudius Gros · 2024
Connects AlphaZero scaling and inverse scaling to the frequency distribution of game states.
Scaling laws for single-agent reinforcement learning
Jacob Hilton, Jie Tang, John Schulman · 2023
Studies how reinforcement-learning performance and optimal model size scale with training compute and environment interactions.
Scaling Laws for Imitation Learning in Single-Agent Games
Jens Tuyls, Dhruv Madeka, Kari Torkkola, et al. · 2023
Studies compute, model, and data scaling for imitation-learning agents in Atari and NetHack.
Training Compute-Optimal Large Language Models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, et al. · 2022
Studies the allocation of model size and training data under a fixed language-model training budget.
Chinchilla Scaling: A replication attempt
Tamay Besiroglu, Ege Erdil, Matthew Barnett, et al. · 2024
Reexamines the parametric fitting procedure and confidence intervals in Chinchilla scaling estimates.
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, et al. · 2020
Measures relationships between language-model loss, model size, dataset size, and training compute.
A Hitchhiker's Guide to Scaling Law Estimation
Leshem Choshen, Yang Zhang, Jacob Andreas · 2024
Studies methods for estimating scaling laws from existing models and intermediate training checkpoints.
Evaluation and uncertainty
Deep Reinforcement Learning at the Edge of the Statistical Precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, et al. · 2021
Studies uncertainty in reinforcement-learning benchmarks and proposes interval estimates and robust aggregate performance measures.
Deep Reinforcement Learning that Matters
Peter Henderson, Riashat Islam, Philip Bachman, et al. · 2017
Studies how experimental choices and random variation affect reproducibility in deep reinforcement learning.
How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments
Cédric Colas, Olivier Sigaud, Pierre-Yves Oudeyer · 2018
Explains how random-seed counts affect statistical power and error rates in reinforcement-learning comparisons.
A Hitchhiker's Guide to Statistical Comparisons of Reinforcement Learning Algorithms
Cédric Colas, Olivier Sigaud, Pierre-Yves Oudeyer · 2019
Compares statistical tests for reinforcement-learning experiments and examines their robustness to violated assumptions.
Empirical Design in Reinforcement Learning
Andrew Patterson, Samuel Neumann, Martha White, et al. · 2023
Provides methods for experimental design, statistical comparison, and control of experimenter bias in reinforcement learning.
AIVAT: A New Variance Reduction Technique for Agent Evaluation in Imperfect Information Games
Neil Burch, Martin Schmid, Matej Moravcik, et al. · 2018
Introduces an unbiased estimator that reduces evaluation variance from chance events and known player strategies.
Re-evaluating Evaluation
David Balduzzi, Karl Tuyls, Julien Perolat, et al. · 2018
Introduces Nash averaging to reduce bias from redundant tasks or agents in evaluation.
Real World Games Look Like Spinning Tops
Wojciech Marian Czarnecki, Gauthier Gidel, Brendan Tracey, et al. · 2020
Studies transitive strength and strategic cycles in games and relates this structure to training with policy populations.
Approximate exploitability: Learning a best response in large games
Finbarr Timbers, Nolan Bard, Edward Lockhart, et al. · 2020
Estimates exploitability in large games by learning a response that targets an agent’s weaknesses.
Training and optimization
What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
Marcin Andrychowicz, Anton Raichuk, Piotr Stańczyk, et al. · 2020
Measures how implementation and design choices affect on-policy reinforcement-learning agents in continuous-control tasks.
Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, et al. · 2020
Studies how implementation details influence the behavior and measured performance of PPO and TRPO.
Batch size-invariance for policy optimization
Jacob Hilton, Karl Cobbe, John Schulman · 2021
Separates proximal and behavior policies to make policy optimization less sensitive to batch size.
An Empirical Model of Large-Batch Training
Sam McCandlish, Jared Kaplan, Dario Amodei, et al. · 2018
Uses gradient noise to estimate useful batch sizes and the tradeoff between compute efficiency and training time.
Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
Greg Yang, Edward J. Hu, Igor Babuschkin, et al. · 2022
Introduces a parametrization that transfers tuned hyperparameters from smaller neural networks to larger ones.
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
Alexander Hägele, Elie Bakouch, Atli Kosson, et al. · 2024
Studies constant learning rates with cooldowns as a way to reuse training runs when estimating scaling laws.
Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, et al. · 2016
Introduces adaptive allocation of training resources to randomly sampled hyperparameter configurations.
A System for Massively Parallel Hyperparameter Tuning
Liam Li, Kevin Jamieson, Afshin Rostamizadeh, et al. · 2018
Introduces asynchronous successive halving for parallel hyperparameter tuning with early stopping.
Understanding Short-Horizon Bias in Stochastic Meta-Optimization
Yuhuai Wu, Mengye Ren, Renjie Liao, et al. · 2018
Studies how short optimization horizons bias hyperparameter selection toward small learning rates.
AutoRL Hyperparameter Landscapes
Aditya Mohan, Carolin Benjamins, Konrad Wienecke, et al. · 2023
Measures how reinforcement-learning hyperparameter landscapes change during training.
Population Based Training of Neural Networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, et al. · 2017
Jointly optimizes neural-network weights and hyperparameter schedules through a population of training runs.
Provably Efficient Online Hyperparameter Optimization with Population-Based Bandits
Jack Parker-Holder, Vu Nguyen, Stephen Roberts · 2020
Uses probabilistic models to guide population-based hyperparameter optimization with fewer parallel agents.
Self-play systems
Emergent Complexity via Multi-Agent Competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, et al. · 2017
Shows how competitive self-play can produce complex behavior and a natural learning curriculum.
Deep Reinforcement Learning from Self-Play in Imperfect-Information Games
Johannes Heinrich, David Silver · 2016
Introduces Neural Fictitious Self-Play for learning approximate equilibria in imperfect-information games.
A Unified Game-Theoretic Approach to Multiagent Reinforcement Learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, et al. · 2017
Introduces a game-theoretic learning framework that trains responses to mixtures of policies and computes policy-selection strategies.
Adversarial Policies Beat Superhuman Go AIs
Tony T. Wang, Adam Gleave, Tom Tseng, et al. · 2022
Trains adversarial Go policies that exploit weaknesses in strong Go agents.
Dota 2 with Large Scale Deep Reinforcement Learning
OpenAI: Christopher Berner, Greg Brockman, Brooke Chan, et al. · 2019
Describes OpenAI Five and the distributed self-play training system used for competitive Dota 2.
Accelerating Self-Play Learning in Go
David J. Wu · 2019
Introduces improvements that accelerate neural-network-guided self-play learning in Go.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, et al. · 2019
Describes AlphaStar’s use of human games and a league of adapting strategies to learn StarCraft II.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, et al. · 2017
Presents a Go agent that learns through self-play reinforcement learning without human game data.
Outracing champion Gran Turismo drivers with deep reinforcement learning
Peter R. Wurman, Samuel Barrett, Kenta Kawamoto, et al. · 2022
Trains Gran Turismo racing agents through deep reinforcement learning with mixed scenarios and rewards for racing conduct.
Model growth and reuse
Net2Net: Accelerating Learning via Knowledge Transfer
Tianqi Chen, Ian Goodfellow, Jonathon Shlens · 2015
Introduces function-preserving transformations that transfer knowledge into deeper or wider neural networks.
Staged Training for Transformer Language Models
Sheng Shen, Pete Walsh, Kurt Keutzer, et al. · 2022
Studies staged growth of transformer language models while preserving loss and training dynamics.
bert2BERT: Towards Reusable Pretrained Language Models
Cheng Chen, Yichun Yin, Lifeng Shang, et al. · 2021
Transfers parameters from smaller pretrained language models to initialize larger models and reduce training cost.
Learning to Grow Pretrained Models for Efficient Transformer Training
Peihao Wang, Rameswar Panda, Lucas Torroba Hennigen, et al. · 2023
Learns linear growth operators that initialize larger transformers from smaller pretrained models.
Kickstarting Deep Reinforcement Learning
Simon Schmitt, Jonathan J. Hudson, Augustin Zidek, et al. · 2018
Uses trained teacher policies to accelerate reinforcement-learning students while allowing students to exceed their teachers.
Reincarnating Reinforcement Learning: Reusing Prior Computation to Accelerate Progress
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, et al. · 2022
Studies how existing agents and prior computation can accelerate the development of new reinforcement-learning agents.
On Warm-Starting Neural Network Training
Jordan T. Ash, Ryan P. Adams · 2019
Studies generalization problems in warm-started neural networks and methods that reduce those problems.
The Primacy Bias in Deep Reinforcement Learning
Evgenii Nikishin, Max Schwarzer, Pierluca D'Oro, et al. · 2022
Studies overfitting to early reinforcement-learning experience and tests partial network resets as a remedy.
Understanding plasticity in neural networks
Clare Lyle, Zeyu Zheng, Evgenii Nikishin, et al. · 2023
Studies why neural networks lose their ability to adapt and tests design choices that preserve plasticity.