cs.MAOct 8, 2026

ConventionPlay: Capability-Limited Training for Robust Ad-Hoc Collaboration

Authors: Abhishek Sriraman, Eleni Vasilaki, Robert Loftin

Organizations: Department of Computer Science Sheffield, UK · Pittsburgh, PA 15213, USA

Abstract

Ad-hoc collaboration often requires agents to identify and adhere to some shared convention within a cooperative task. Existing work on reinforcement learning (RL) for ad-hoc collaboration focuses on training agents that adapt to the conventions established by their partners. These methods fail to consider the possibility that while some partners might follow only a single fixed convention, others may themselves be capable of adapting to multiple conventions. Here we present ConventionPlay, an RL-based approach that teaches agents to discover their partner's optimal convention by training against a learned population of partners that exhibit different degrees of adaptability across conventions. Some of these partners follow a single, fixed convention, while others are able to adapt to a subset of the possible conventions for the task in question. The existence of partners that support a limited subset of conventions forces agents trained against this population to actively probe their partner's capabilities, and steer their partner towards the most effective joint strategy that they are capable of following. Our experimental results demonstrate that agents trained via ConventionPlay achieve superior performance to existing ad-hoc collaboration methods against test populations of partners that are compatible with multiple conventions.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ConventionPlay: Capability-Limited Training for Robust Ad-Hoc Collaboration

    Apr 20, 2026Abhishek Sriraman, Eleni Vasilaki, Robert LoftinMulti-Agent CoordinationMulti-Agent Collaboration

  2. Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

    Jul 29, 2026Peter Tisnikar, Maja Swieczkowska, Benteng Ma +2Bayesian InferenceAd Hoc Teamwork

  3. ROTATE: Regret-driven Open-ended Training for Ad Hoc Teamwork

    May 29, 2025Caroline Wang, Arrasy Rahman, Benjamin Nativi +4Multi-Agent CollaborationAd Hoc Teamwork