Joint attention - the human ability to share a common visual or cognitive focus with others - enables a meeting of minds that lets us coordinate even with unfamiliar partners. In this work we investigate whether equipping AI agents with a similar mechanism can enable such zero-shot coordination. We introduce Mutual Attention for zero-shot TEaming (MATE): a novel multi-agent reinforcement learning method inspired by human joint attention. MATE encourages agents to coordinate their actions by aligning their visual attention on scene-salient objects during the interaction rather than relying on arbitrary partner-dependent conventions established during training. Unlike symmetry-breaking approaches that merely prevent brittle conventions from emerging, MATE actively promotes coordination through an environment-grounded signal that is naturally shared across partners. We evaluate MATE on three benchmarks: our Card Alignment Game, designed to isolate brittle convention formation, and the more challenging Level-Based Foraging and OvercookedV2 benchmarks. Our experiments consistently show that a joint-attention-inspired signal improves coordination with unknown partners, underlining MATE's potential as a general coordination mechanism that complements and surpasses symmetry-breaking approaches.
Figures & tables
Figure 1: Zero-Shot Coordination using joint attention. Alice and Bob , trained independently, must pick the same card without explicit communication. Self-play agents commit to arbitrary conventions, other-play agents have no signal at all to coordinate on, while MATE agents read each other’s attention and converge on a shared choice.
Figure 2: Sample episode in the Card Alignment Game. Alice and Bob must select the same card in the unpermuted layout (left). To prevent position or colour conventions, they observe differently permuted frames. Finally, both agents choose independently.
Figure 3: SP (hatched) and XP (solid) scores for the Card Alignment Game (a), LBF (b) and OvercookedV2 (c). Error bars show SEM. MATE achieves the highest XP across all three environments.
Figure 4: Card Alignment Game attention maps. (a) and (b) show per-card attention in the unpermuted frame. MATE progressively aligns attention between agents and selects the same card, whereas OP maintains divergent attention and chooses different cards. (c) shows card-attention JSD over time; lower values indicate stronger alignment.
Figure 5: Level-Based Foraging attention maps. (a) and (b) show the combined attention of two independently trained agents in the unpermuted frame. The colour scale indicates normalised attention intensity. When several apples are equally valid targets, MATE agents converge on the same apple, whereas OP agents attend to different regions and fail to establish a common target.
Figure 6: OvercookedV2 attention maps from two moments of interaction between MATE agents. Red and blue boxes delimit the agents’ respective fields of view; contour lines show their combined attention in the unpermuted frame.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Scores of HE IPPO for SP and XP, against the entropy coefficient. Shaded regions show SEM.
Hyperparameter
Card Align.
LBF
OCv2
LR
{3,5,8}×10−4
4×10−4
{2.5,4}×10−4
LR annealing
False
True
True
Steps
5×106
3×106
3×107
Parallel envs
{512,1024}
{128,256}
{192,256}
Rollout
8
128
256
Epochs
4
Appendix
Table 1: Training hyperparameters for Card Alignment Game, LBF, and OvercookedV2. Braces show the values considered during tuning; bold values were selected. Empty entries inherit Card Alignment Game values.