Self-Play RL

RL: Reinforcement Learning

Momentum

6 papers in the last four weeks, up 50% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 53

All topics
CardsList
  1. Faynt: Scaling and Optimizing Policies for Competitive Melee

    Oct 1, 2026Ali Janati, Nikita Kuzmin, Rohit Swamy +1Reinforcement LearningSelf-Play RL

  2. Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions

    Sep 29, 2026Ondřej Kubíček, Viliam Lisý, Tuomas SandholmGame-Playing AgentsSelf-Play RL

  3. Learning to Optimize through Solver-Grounded Self-Play

    Sep 28, 2026Xia Jiang, Yaoxin Wu, Chenyu Zhou +3Self-Play RLNatural Language to Optimization Modeling

  4. The Surprising Effectiveness of Approximate Value Iteration in Self-Play

    Sep 8, 2026Raphael Boige, Amine Boumaza, Bruno ScherrerGame-Playing AgentsSelf-Play RL

  5. SQL-Zero: Self-Evolving Text-to-SQL

    Sep 7, 2026Daniel Machado Pedrozo, Julia Soares Dollis, Bryan Lincoln Marques de Oliveira +3Reinforcement LearningSelf-Play RL

  6. Local Updates, Global Learning (LUGL): Playing Games with non-incremental Learners

    Sep 3, 2026David Milec, Spyridon Samothrakis, Michael Fairbank +1Gradient-Boosted Decision TreesGame-Playing Agents

  7. What Emerges and What Breaks in Self-Play Driving

    Aug 31, 2026Laur Sisask, Ardi Tampuu, Tambet MatiisenReward HackingSelf-Play RL

  8. SPADE: Self-Play in Adaptive Synthetic Executable Environments

    Aug 19, 2026Bo Liu, Simon Yu, Yiding Jiang +15Self-Play RLLLM Agent Self-Improvement

  9. Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

    Jul 31, 2026Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai +6Self-Play RLAgent Skill Learning

  10. Pictura: Perspective-View Self-Play at Scale for Driving

    Jul 28, 2026Yuan Yin, Elias Ramzi, Marc Lafon +8Self-Play RLRL for Autonomous Driving

  11. From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

    Jul 26, 2026Qinsi Wang, Jing Shi, Huazheng Wang +8Self-Play RLReinforcement Learning with Verifiable Rewards

  12. TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

    Jul 14, 2026Zhouchonghao Wu, Akshay Rangesh, Weixin Li +4Self-Play RLRL for Autonomous Driving

  13. World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning

    Jul 12, 2026Tong Nie, Yuewen Mei, Junlin He +3Adversarial TrainingSelf-Play RL

  14. AlphaZero in Sparsely Rewarded Games: Limits and Auxiliary Supervision

    Jul 9, 2026Brent Kong, Tejas Ram, Tony Yue YuGame-Playing AgentsSelf-Play RL

  15. A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong

    Jul 7, 2026Nima Kelidari, Mohammadsaeed Haghi, Mahdi SalmaniReinforcement LearningGame-Playing Agents

  16. Anchored Self-Play for Code Repair

    Jul 3, 2026Caroline Choi, Zeyneb Kaya, Shirley Wu +3RL for Code GenerationAutomated Program Repair

  17. EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games

    Jun 22, 2026Tristan Maidment, JB Lanier, Chase McDonald +5Self-Play RLImperfect-Information Games

  18. Superhuman AI for Generals.io Using Self-Play Reinforcement Learning

    Jun 22, 2026Matej Straka, Viliam Lisý, Martin SchmidGame-Playing AgentsSelf-Play RL

  19. Scaling Self-Play for End-to-End Driving

    Jun 17, 2026Luke Rowe, Roger Girgis, Rodrigue de Schaetzen +6Self-Play RLEnd-to-End Autonomous Driving

  20. Human-like autonomy emerges from self-play and a pinch of human data

    Jun 11, 2026Daphne Cornelisse, Julian Hunt, Zixu Zhang +4Self-Play RLAutonomous Driving

  21. CoPark: Learning Reactive Parking via Self-Play

    Jun 2, 2026Jiarong Wei, Yanxing Chen, Sinuo Song +3Multi-Agent CoordinationSelf-Play RL

  22. S-SPPO: Semantic-Calibrated Self-Play Preference Optimization

    Jun 1, 2026Xiwen Chen, Wenhui Zhu, Jingjing Wang +13LLM AlignmentSelf-Play RL

  23. SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

    May 29, 2026Wai-Chung Kwan, Aryo Pradipta Gema, Joshua Ong Jun Leang +1Open-Ended GenerationSelf-Play RL

  24. Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL

    May 21, 2026Sophia Xiao Pu, Zhaotian Weng, Chengzhi Liu +4Reinforcement LearningSelf-Play RL