Self-Play RL

RL: Reinforcement Learning

Momentum

6 papers in the last four weeks, up 50% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 53

All topics
CardsList
  1. Faynt: Scaling and Optimizing Policies for Competitive Melee

    Oct 1, 2026Ali Janati, Nikita Kuzmin, Rohit Swamy +1Reinforcement LearningSelf-Play RL

  2. Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions

    Sep 29, 2026Ondřej Kubíček, Viliam Lisý, Tuomas SandholmGame-Playing AgentsSelf-Play RL

  3. Learning to Optimize through Solver-Grounded Self-Play

    Sep 28, 2026Xia Jiang, Yaoxin Wu, Chenyu Zhou +3Self-Play RLNatural Language to Optimization Modeling

  4. The Surprising Effectiveness of Approximate Value Iteration in Self-Play

    Sep 8, 2026Raphael Boige, Amine Boumaza, Bruno ScherrerGame-Playing AgentsSelf-Play RL

  5. SQL-Zero: Self-Evolving Text-to-SQL

    Sep 7, 2026Daniel Machado Pedrozo, Julia Soares Dollis, Bryan Lincoln Marques de Oliveira +3Reinforcement LearningSelf-Play RL

  6. Local Updates, Global Learning (LUGL): Playing Games with non-incremental Learners

    Sep 3, 2026David Milec, Spyridon Samothrakis, Michael Fairbank +1Gradient-Boosted Decision TreesGame-Playing Agents

  7. What Emerges and What Breaks in Self-Play Driving

    Aug 31, 2026Laur Sisask, Ardi Tampuu, Tambet MatiisenReward HackingSelf-Play RL

  8. SPADE: Self-Play in Adaptive Synthetic Executable Environments

    Aug 19, 2026Bo Liu, Simon Yu, Yiding Jiang +15Self-Play RLLLM Agent Self-Improvement

  9. Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

    Jul 31, 2026Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai +6Self-Play RLAgent Skill Learning

  10. Pictura: Perspective-View Self-Play at Scale for Driving

    Jul 28, 2026Yuan Yin, Elias Ramzi, Marc Lafon +8Self-Play RLRL for Autonomous Driving

  11. From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

    Jul 26, 2026Qinsi Wang, Jing Shi, Huazheng Wang +8Self-Play RLReinforcement Learning with Verifiable Rewards

  12. TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

    Jul 14, 2026Zhouchonghao Wu, Akshay Rangesh, Weixin Li +4Self-Play RLRL for Autonomous Driving

  13. World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning

    Jul 12, 2026Tong Nie, Yuewen Mei, Junlin He +3Adversarial TrainingSelf-Play RL

  14. AlphaZero in Sparsely Rewarded Games: Limits and Auxiliary Supervision

    Jul 9, 2026Brent Kong, Tejas Ram, Tony Yue YuGame-Playing AgentsSelf-Play RL

  15. A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong

    Jul 7, 2026Nima Kelidari, Mohammadsaeed Haghi, Mahdi SalmaniReinforcement LearningGame-Playing Agents

  16. Anchored Self-Play for Code Repair

    Jul 3, 2026Caroline Choi, Zeyneb Kaya, Shirley Wu +3RL for Code GenerationAutomated Program Repair

  17. EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games

    Jun 22, 2026Tristan Maidment, JB Lanier, Chase McDonald +5Self-Play RLImperfect-Information Games

  18. Superhuman AI for Generals.io Using Self-Play Reinforcement Learning

    Jun 22, 2026Matej Straka, Viliam Lisý, Martin SchmidGame-Playing AgentsSelf-Play RL

  19. Scaling Self-Play for End-to-End Driving

    Jun 17, 2026Luke Rowe, Roger Girgis, Rodrigue de Schaetzen +6Self-Play RLEnd-to-End Autonomous Driving

  20. Human-like autonomy emerges from self-play and a pinch of human data

    Jun 11, 2026Daphne Cornelisse, Julian Hunt, Zixu Zhang +4Self-Play RLAutonomous Driving

  21. CoPark: Learning Reactive Parking via Self-Play

    Jun 2, 2026Jiarong Wei, Yanxing Chen, Sinuo Song +3Multi-Agent CoordinationSelf-Play RL

  22. S-SPPO: Semantic-Calibrated Self-Play Preference Optimization

    Jun 1, 2026Xiwen Chen, Wenhui Zhu, Jingjing Wang +13LLM AlignmentSelf-Play RL

  23. SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

    May 29, 2026Wai-Chung Kwan, Aryo Pradipta Gema, Joshua Ong Jun Leang +1Open-Ended GenerationSelf-Play RL

  24. Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL

    May 21, 2026Sophia Xiao Pu, Zhaotian Weng, Chengzhi Liu +4Reinforcement LearningSelf-Play RL

  25. GeoX: Mastering Geospatial Reasoning Through Self-Play and Verifiable Rewards

    May 19, 2026Kyeongjin Ahn, Seungeon Lee, Krishna P. Gummadi +1Self-Play RLGeospatial Reasoning

  26. Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment

    May 17, 2026Yucong Huang, Xiucheng Li, Kaiqi Zhao +1Reward ModelingLLM Alignment

  27. PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play

    May 16, 2026Roger Creus Castanyer, Geoffrey Bradway, Lorenz Wolf +3Self-Play RLRL for Language Model Reasoning

  28. Data-Augmented Game Starts for Accelerating Self-Play Exploration in Imperfect Information Games

    May 14, 2026JB Lanier, Nathan Monette, Pierre Baldi +1RL BenchmarksSelf-Play RL

  29. Offline Two-Player Zero-Sum Markov Games with KL Regularization

    May 13, 2026Claire Chen, Yuheng Zhang, Xinyu Liu +3Nash EquilibriumReinforcement Learning

  30. Beyond Self-Play and Scale: A Behavior Benchmark for Generalization in Autonomous Driving

    May 11, 2026Aron Distelzweig, Faris Janjoš, Andreas Look +7Autonomous Driving BenchmarksSelf-Play RL

  31. G-Zero: Self-Play for Open-Ended Generation from Zero Data

    May 11, 2026Chengsong Huang, Haolin Liu, Tong Zheng +7Open-Ended GenerationSelf-Play RL

  32. Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs

    May 11, 2026Wu Li, Yigeng Zhou, Zesheng Shi +3LLM AlignmentSelf-Play RL

  33. Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents

    May 9, 2026Minzheng Wang, Run Luo, Yanbo Wang +6Reinforcement LearningSelf-Play RL

  34. The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play

    May 8, 2026Gabriele La Malfa, Emanuele La Malfa, Saar Cohen +4Self-Play RLLanguage Model Safety Evaluation

  35. SPARK: Self-Play with Asymmetric Reward from Knowledge Graphs

    May 7, 2026Hyobin Park, Taeseop Kim, Dong-Geol ChoiAI for ScienceSelf-Play RL

  36. Scaling Self-Play with Self-Guidance

    Apr 22, 2026Luke Bailey, Kaiyue Wen, Kefan Dong +2Self-Play RLAutomated Theorem Proving

  37. Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text

    Apr 21, 2026Chengyu Huang, Sheng-Yen Chou, Zhengxin Zhang +1Self-Play RLLLM Post-Training

  38. Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play

    Apr 20, 2026Xiachong Feng, Deyi Yin, Xiaocheng Feng +9Self-Play RLRL for Language Model Reasoning

  39. Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification

    Apr 18, 2026Antonio Valerio Miceli Barone, Poon Tsz NokAdversarial TrainingSelf-Play RL

  40. Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability

    Jan 26, 2026Shobhita Sundaram, John Quan, Ariel Kwiatkowski +3Self-Play RLRL for Language Model Reasoning

  41. Scalable Decision Making for Games of Imperfect Information

    Nov 10, 2025Samuel Sokota, Eugene Vinitsky, Hengyuan Hu +3Game-Playing AgentsInference-Time Search

  42. OpenSIR: Open-Ended Self-Improving Reasoner

    Nov 1, 2025Wai-Chung Kwan, Joshua Ong Jun Leang, Pavlos Vougiouklis +3Self-Play RLRL for Language Model Reasoning

  43. Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

    Jun 9, 2025Mickel Liu, Liwei Jiang, Yancheng Liang +4Self-Play RLAdversarial Attacks on LLMs