Self-Play RL

RL: Reinforcement Learning

Momentum

6 papers in the last four weeks, up 50% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 53

All topics
CardsList
  1. GeoX: Mastering Geospatial Reasoning Through Self-Play and Verifiable Rewards

    May 19, 2026Kyeongjin Ahn, Seungeon Lee, Krishna P. Gummadi +1Self-Play RLGeospatial Reasoning

  2. Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignment

    May 17, 2026Yucong Huang, Xiucheng Li, Kaiqi Zhao +1Reward ModelingLLM Alignment

  3. PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play

    May 16, 2026Roger Creus Castanyer, Geoffrey Bradway, Lorenz Wolf +3Self-Play RLRL for Language Model Reasoning

  4. Data-Augmented Game Starts for Accelerating Self-Play Exploration in Imperfect Information Games

    May 14, 2026JB Lanier, Nathan Monette, Pierre Baldi +1RL BenchmarksSelf-Play RL

  5. Offline Two-Player Zero-Sum Markov Games with KL Regularization

    May 13, 2026Claire Chen, Yuheng Zhang, Xinyu Liu +3Nash EquilibriumReinforcement Learning

  6. Beyond Self-Play and Scale: A Behavior Benchmark for Generalization in Autonomous Driving

    May 11, 2026Aron Distelzweig, Faris Janjoš, Andreas Look +7Autonomous Driving BenchmarksSelf-Play RL

  7. G-Zero: Self-Play for Open-Ended Generation from Zero Data

    May 11, 2026Chengsong Huang, Haolin Liu, Tong Zheng +7Open-Ended GenerationSelf-Play RL

  8. Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs

    May 11, 2026Wu Li, Yigeng Zhou, Zesheng Shi +3LLM AlignmentSelf-Play RL

  9. Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents

    May 9, 2026Minzheng Wang, Run Luo, Yanbo Wang +6Reinforcement LearningSelf-Play RL

  10. The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play

    May 8, 2026Gabriele La Malfa, Emanuele La Malfa, Saar Cohen +4Self-Play RLLanguage Model Safety Evaluation

  11. SPARK: Self-Play with Asymmetric Reward from Knowledge Graphs

    May 7, 2026Hyobin Park, Taeseop Kim, Dong-Geol ChoiAI for ScienceSelf-Play RL

  12. Scaling Self-Play with Self-Guidance

    Apr 22, 2026Luke Bailey, Kaiyue Wen, Kefan Dong +2Self-Play RLAutomated Theorem Proving

  13. Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text

    Apr 21, 2026Chengyu Huang, Sheng-Yen Chou, Zhengxin Zhang +1Self-Play RLLLM Post-Training

  14. Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play

    Apr 20, 2026Xiachong Feng, Deyi Yin, Xiaocheng Feng +9Self-Play RLRL for Language Model Reasoning

  15. Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification

    Apr 18, 2026Antonio Valerio Miceli Barone, Poon Tsz NokAdversarial TrainingSelf-Play RL

  16. Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability

    Jan 26, 2026Shobhita Sundaram, John Quan, Ariel Kwiatkowski +3Self-Play RLRL for Language Model Reasoning

  17. Scalable Decision Making for Games of Imperfect Information

    Nov 10, 2025Samuel Sokota, Eugene Vinitsky, Hengyuan Hu +3Game-Playing AgentsInference-Time Search

  18. OpenSIR: Open-Ended Self-Improving Reasoner

    Nov 1, 2025Wai-Chung Kwan, Joshua Ong Jun Leang, Pavlos Vougiouklis +3Self-Play RLRL for Language Model Reasoning

  19. Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

    Jun 9, 2025Mickel Liu, Liwei Jiang, Yancheng Liang +4Self-Play RLAdversarial Attacks on LLMs