Self-Play

Latest papers 98

All topics
CardsList
  1. Learning Explainable Representations of Complex Game-playing Strategies

    Oct 6, 2026Abhijeet Krishnan, Colin M. Potts, Arnav Jhala +3Self-Play

  2. Game-Guided Skill Discovery through Self-Play for Playable Agent Control

    Sep 30, 2026Seungeun Rho, Jeonghwan Kim, Xue Bin Peng +1Unsupervised Skill DiscoveryHuman-To-Robot Transfer

  3. Teaching LLMs to Generate Challenging MILP Instances via Solver Feedback

    Sep 29, 2026Jitin Singla, Parikshit Pareek, Pratik Jawanpuria +1Mixed-Integer ProgrammingHard

  4. Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions

    Sep 29, 2026Ondřej Kubíček, Viliam Lisý, Tuomas SandholmMean Field GamesMirror Descent

  5. Learning to Optimize through Solver-Grounded Self-Play

    Sep 28, 2026Xia Jiang, Yaoxin Wu, Chenyu Zhou +3Optimization ModelingSelf-Play

  6. Self-Play Pretraining with Zero Data

    Sep 24, 2026Aditya Cowsik, Kfir Dolev, Michael Y. Li +4PretrainingSelf-Play

  7. The Surprising Effectiveness of Approximate Value Iteration in Self-Play

    Sep 8, 2026Raphael Boige, Amine Boumaza, Bruno ScherrerSelf-PlayMonte Carlo Tree Search

  8. SQL-Zero: Self-Evolving Text-to-SQL

    Sep 7, 2026Daniel Machado Pedrozo, Julia Soares Dollis, Bryan Lincoln Marques de Oliveira +3Text-To-SqlSelf-Play

  9. Local Updates, Global Learning (LUGL): Playing Games with non-incremental Learners

    Sep 3, 2026David Milec, Spyridon Samothrakis, Michael Fairbank +1Singular Learning TheoryImperfect-Information Games

  10. What Emerges and What Breaks in Self-Play Driving

    Aug 31, 2026Laur Sisask, Ardi Tampuu, Tambet MatiisenSelf-PlayHuman Driving

  11. Beyond Uncertainty: Multi-Solver Disagreement Rewards for Self-Evolving Reasoning Curricula

    Aug 30, 2026Vinoth Selvendran, Zhanming ZhangSelf-PlayCurriculum

  12. SPADE: Self-Play in Adaptive Synthetic Executable Environments

    Aug 19, 2026Bo Liu, Simon Yu, Yiding Jiang +15Synthetic EnvironmentsSelf-Play

  13. Do LLMs Beat Nash? Testing Decentralized Coordination in Self-Play Multi-Agent Games

    Aug 12, 2026Deborah Sinishaw, Qile Zhu, Edwin Meriaux +1Multi-Agent Large Language Model SystemsMulti-Agent Coordination

  14. Solver-Guided Reasoning for Mixed-Equilibrium Strategies

    Aug 7, 2026Han Wang, Philippe Beardsell, Boning Li +4PokerNash Equilibrium

  15. SearchMaster: Grounded and Regulated Self-Play for Search Agents

    Aug 3, 2026Wentao Tan, Qiong Cao, Jiaqi Wang +1Search AgentsSelf-Play

  16. Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

    Jul 31, 2026Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai +6Self-Evolving AgentsSkill Evolution

  17. Pictura: Perspective-View Self-Play at Scale for Driving

    Jul 28, 2026Yuan Yin, Elias Ramzi, Marc Lafon +8Waymo Open Motion DatasetSelf-Play

  18. GPT-Red: Automated Red Teaming via Self-Play at Scale

    Jul 28, 2026Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15Red-TeamingAttacker Large Language Model

  19. Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

    Jul 24, 2026Siyuan Huang, Pengyu Cheng, Haotian Liu +10Self-EvolutionSelf-Play

  20. Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

    Jul 22, 2026Siqian Tong, Xuan Li, Chaozhuo Li +5Audio UnderstandingAudio Understanding

  21. TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

    Jul 14, 2026Zhouchonghao Wu, Akshay Rangesh, Weixin Li +4Waymo Open Motion DatasetSelf-Play

  22. World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning

    Jul 12, 2026Tong Nie, Yuewen Mei, Junlin He +3World Model PlanningMotion Planning

  23. AlphaZero in Sparsely Rewarded Games: Limits and Auxiliary Supervision

    Jul 9, 2026Brent Kong, Tejas Ram, Tony Yue YuSelf-PlayChess

  24. A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong

    Jul 7, 2026Nima Kelidari, Mohammadsaeed Haghi, Mahdi SalmaniImperfect-Information GamesSelf-Play

  25. FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games

    Jul 7, 2026Chase McDonald, Nathan Tsang, Wesley N. KerrImperfect-Information GamesSkillgym

  26. Predicting Drafted Deck Strength for "Magic: the Gathering"

    Jul 6, 2026Tomas Rigaux, Hisashi KashimaCardSelf-Play

  27. Anchored Self-Play for Code Repair

    Jul 3, 2026Caroline Choi, Zeyneb Kaya, Shirley Wu +3Automated Program RepairRepair

  28. Parametric Open Source Games

    Jun 25, 2026Aleksandar Todorov, Jesse ten Napel, Alexander MüllerGame TheoryNash Equilibrium

  29. EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games

    Jun 22, 2026Tristan Maidment, JB Lanier, Chase McDonald +5Self-PlayImperfect-Information Games

  30. Superhuman AI for Generals.io Using Self-Play Reinforcement Learning

    Jun 22, 2026Matej Straka, Viliam Lisý, Martin SchmidSelf-PlaySimulation-Based Reinforcement Learning

  31. GeoRouteNet: Geometry-Enhanced Non-Autoregressive Neural Solver for the Traveling Salesman Problem

    Jun 22, 2026Xiang LiSalesman ProblemNeural Solvers

  32. Scaling Self-Play for End-to-End Driving

    Jun 17, 2026Luke Rowe, Roger Girgis, Rodrigue de Schaetzen +6Causality-Aware End-To-End Autonomous DrivingSelf-Play

  33. TerraTransfer: Learning End-to-End Driving Policies Without Expert Demonstrations

    Jun 16, 2026Zikang Xiong, Weixin Li, Zhouchonghao Wu +6Autonomous DrivingUnsupervised On-Policy Self-Distillation

  34. Discovering Lattice Reduction Strategies via Self-Play

    Jun 13, 2026Mohamed Malhou, Kristin Lauter, Ludovic PerretSelf-PlayMonte Carlo Tree Search

  35. Human-like autonomy emerges from self-play and a pinch of human data

    Jun 11, 2026Daphne Cornelisse, Julian Hunt, Zixu Zhang +4Self-PlayRandom

  36. CoPark: Learning Reactive Parking via Self-Play

    Jun 2, 2026Jiarong Wei, Yanxing Chen, Sinuo Song +3Residual PolicyMulti-Agent Reinforcement Learning

  37. Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs

    Jun 2, 2026Wenqi Chen, Ziyan Zhang, Bin Wang +3Vulnerable CodeLarge Language Model Safety

  38. S-SPPO: Semantic-Calibrated Self-Play Preference Optimization

    Jun 1, 2026Xiwen Chen, Wenhui Zhu, Jingjing Wang +13Self-Play

  39. SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

    May 29, 2026Wai-Chung Kwan, Aryo Pradipta Gema, Joshua Ong Jun Leang +1Self-PlayTask-Specific Rubrics

  40. Not All Synthetic Data Is Yours to Learn From

    May 29, 2026Sina Alemohammad, Li Chen, Richard G. Baraniuk +1Synthetic Training DataSynthetic Data

  41. On the Geometry of Games and their Solvers

    May 28, 2026Yaqi Sun, Julian Ma, David MguniGame TheoryEquilibrium

  42. Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning

    May 21, 2026Ismail Geles, Leonard Bauersfeld, Markus Wulfmeier +1Multi-Agent Reinforcement LearningMulti-Robot Systems

  43. Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL

    May 21, 2026Sophia Xiao Pu, Zhaotian Weng, Chengzhi Liu +4Self-PlayUnsupervised On-Policy Self-Distillation

  44. GeoX: Mastering Geospatial Reasoning Through Self-Play and Verifiable Rewards

    May 19, 2026Kyeongjin Ahn, Seungeon Lee, Krishna P. Gummadi +1Geospatial ReasoningGeogs-Slam