Zero-shot coordination in embodied settings requires acting while the partner is intermittently out of view, leaving existing methods with ambiguous partner representations and uncertainty over hidden partner states. We propose Predicting Intention of Partner (PIP) to jointly address these challenges. PIP uses a Joint-view VAE to distill richer training-time evidence from the union of both agents' local observations into a partner representation available from local observations alone. Partner-state Belief networks further infer the partner's hidden location and behavioral tendencies from the ego agent's interaction history. We evaluate PIP in Burrito-PO, Overcooked-PO, and a Melting Pot substrate, together with a human evaluation in Burrito-PO. PIP attains the highest mean performance among the compared methods across all three benchmarks. Human evaluation and diagnostic analyses further support coordination with unseen partners and the contributions of both components under partner occlusion.
Figures & tables
Figure 1 : Illustration of coordination under dynamic and spatial partial observability in Burrito-PO. (a) Under full observation, the partner state is directly available. (b) Under a restricted local view, the partner may leave the visible region, creating ambiguity over its location and behavior.
Method
Open
Hallway
FC
Ring
FCP
141.6 ± 9.6
57.1 ± 2.7
21.9 ± 6.0
48.3 ± 6.1
MEP
115.9 ± 10.3
58.2 ± 2.6
17.6 ± 5.8
51.1 ± 5.4
E3T
55.3 ± 3.6
44.5 ± 5.7
21.0 ± 4.1
27.5 ± 3.1
ERS
101.0 ± 7.8
36.6 ± 4.5
24.0 ± 4.3
48.9 ± 5.5
GAMMA
163.3 ± 10.6
57.7 ± 3.3
19.8 ± 5.3
48.5 ± 5.6
GOAT
165.5 ± 10.6
61.5 ± 5.1
19.0 ± 5.3
77.3 ± 5.9
Table 1: Evaluation with 12 held-out behavior-preference partners unseen during training. The best result in each layout is shown in bold . Mean ± standard error.
Table 3
Figure 2 : Human evaluation in the Burrito-PO environment. Participants interacted with GAMMA, GOAT, and our method. Results report (a) sparse reward, (b) cooperation preference ranking, and (c) post-interaction survey scores on coordination quality.
Figure 5Figure 6
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Preference
Reward Shape
Plating Ingredients
−30,20
Washing Plates
−30,20
Delivering Dishes
−30,20
Chopping Ingredients
−30,20
Potting Rice
−30,20
Grilling Meat/Mushroom
−30,20
Appendix
Table B.1 : Event-based BP features and corresponding reward design in Burrito-PO.
Events
Weights
Put an onion / dish / soup / tomato
0
onto the counter
Pickup an onion / dish / soup / tomato
0
from the counter
Pickup an onion from the onion dispenser
−20,0,10
Pickup an tomato from the tomato dispenser
−20,0,10
Appendix
Table B.2: Event-based BP features and corresponding reward design in Overcooked-PO
Figure C.1 : Partner-visibility ratio (%) by view size and layout, measured from Self-Play rollouts using the MEP algorithm.
Hyperparameter
Burrito-PO
Overcooked-PO
Melting Pot
CNN kernels
[3, 3], [3, 3], [3, 3]
[3, 3], [3, 3], [3, 3]
[8, 8], [4, 4], [3, 3]
CNN channels
[32, 64, 32]
hidden layer size
[64]
recurrent layer size
64
activation function
ReLU
weight decay
0
Appendix
Table C.1 : Hyperparameters for policy models in each environment.
Hyperparameter
Value
Behavioral Belief
λbehavior
0.1
Koff
10
γrec
0.90
γten
0.97
αrec
1.0
Appendix
Table C.2 : Partner-state Belief hyperparameters for target policy models.
Metric
Baseline
PIP (Ours)
Ratio
Total parameters
463K
591K
1.28 ×
Actor rollout inference time
1.13 ms
2.05 ms
1.81 ×
Actor rollout FLOPs
0.381G
0.764G
2.01 ×
Appendix
Table C.3: Representative computational costs at the training rollout batch size of 64. Baseline corresponds to GAMMA / GOAT, whose target policies share an identical architecture.
hyperparameter
Stage 1 (teacher pretrain)
Stage 2 (student distill)
trained module
teacher (enc + dec)
student (enc + dec), teacher frozen
encoder input
(o1∪o2)⊕opartner
opartner
λKD
–
10.0
epoch
500
500
shared across stages
hyperparameter
Burrito-PO
Overcooked-PO
Melting Pot
Appendix
Table C.4 : Hyperparameters for the Joint-view VAE.
Figure E.1 : Burrito environment layouts used in our experiments. Each layout induces different coordination patterns under ego-centric partial observability.
Figure E.2 : The Overcooked environment layout used in our experiments, counter circuit.
Figure E.3 : The Melting Pot coop mining maps used in our experiments. Grey walls occlude the agents’ view, iron ore (grey) gives +1 when mined alone, whereas gold ore (yellow) gives +8 to each agent only if both mine the same ore within a 3-step window.
Figure F.1 : t-SNE visualization of episode-level posterior means from each encoder on held-out trajectories from 12 BP partners (144 episodes per partner; perplexity 50, PCA initialization). Each point corresponds to one episode and is colored by the partner policy that generated it, with the same color used for the same partner across panels. Compared with (a) Fixed σ and (b) Het σ , (c) Ours produces more clearly separated partner-specific latent clusters.
Figure F.2 : Partner-state Belief analysis. (a) Current Location Belief accuracy by Time Since Last Seen (TSLS). (b) Δ -step future Location Belief accuracy. (c) Δ -step future Behavioral Belief accuracy.
Method
Score
GAMMA
143.3 ± 11.2
GOAT
103.0 ± 8.6
PIP
256.2 ± 15.1
Appendix
Table F.1: Full-observation coordination performance on the Burrito Open layout with 12 held-out behavior-preference partners.
Method
Cap.
Sup.
Reward
GAMMA
No
No
163.3 ± 10.6
+ Capacity control
Yes
No
145.8 ± 5.7
+ Supervision control
No
Yes
88.7 ± 10.2
+ Joint control
Yes
Yes
96.0 ± 7.1
PIP (Ours)
Yes
Yes
249.0 ± 12.2
Appendix
Table F.2: Capacity and supervision controls on Burrito-PO Open. Cap. denotes capacity approximately matched to PIP, and Sup. denotes direct auxiliary supervision using privileged partner-state labels.
Figure F.3 : Return of PIP against the M=12 partners as the corruption rate r increases from 0 to 0.2 (mean ± s.e.), under sparse (blackout) and noisy (per-channel shuffle) observation corruption. The dashed line is GAMMA with clean observations ( r=0 ).
Figure G.1 : Participant consent form for the Burrito human-AI coordination study, outlining the study purpose, procedure, anonymous data collection, exclusion of personal identifiers, data protection policy, potential risks, voluntary participation, and withdrawal conditions.
#
Step
Instruction
1
Movement
Use the arrow keys to move in all four directions. Participants were informed that the game only resumes after pressing Enter or the yellow “Continue” button.
2
Read menu
Inspect the order cards at the bottom of the screen. Each card indicates the requested burrito type and the remaining ticks before the order expires. Orders can be delivered in any sequence.
3
Pick meat
Move to the meat dispenser and press Space to pick up a piece of meat.
4
Chop meat
Place the meat on the cutting board, press Space repeatedly to chop it, and pick up the resulting chopped meat.
5
Grill meat
Place the chopped meat on the grill to cook it.
6
Boil rice
Pick up rice from the rice dispenser and place it in the pot to boil.
Appendix
Table G.1 : Tutorial steps provided before the human-AI coordination survey. The tutorial familiarized participants with the game controls, object interactions, burrito recipes, and failure-recovery mechanics.
Figure G.2 : Tutorial and survey interfaces. The left panel shows step-by-step instructions, object icons, and the ego-centric partial-observation view. The right panel shows the main-survey gameplay interface, where participants cooperate with an anonymized AI partner under ego-centric partial observability.
Figure G.3 : Post-game evaluation materials. The left panel shows the post-episode evaluation survey assessing each AI partner after gameplay. The right panel shows the final preference ranking of the three AI partners based on collaboration experience.
Metric
GAMMA
GOAT
PIP (Ours)
vs. GAMMA
vs. GOAT
Sparse reward mean ± SE
355.7±19.7
370.4±18.4
476.0±22.7
+33.8%
+28.5%
Top-1 preference
10/66 , 15.2%
8/66 , 12.1%
48/66 , 72.7%
+57.5 pts.
+60.6 pts.
Appendix
Table G.2 : Human evaluation performance and preference summary ( N=66 participants).
Outcome
Friedman test
PIP vs. GAMMA
PIP vs. GOAT
χ2(2)
p
p
r
p
r
Sparse reward
34.95
2.57×10−8
2.43×10−7
+0.74
3.16×10−6
+0.66
Preference rank
37.85
6.04×10−9
4.80×10−6
+0.62
1.16×10−6
+0.67
Appendix
Table G.3 : Within-subject statistical tests for task performance and preference ranking.
Category
GAMMA
GOAT
PIP
vs. GAMMA
vs. GOAT
Coordination
3.68
3.26
5.01
+36.1%
+53.7%
Predictability
3.33
3.11
4.83
+45.0%
+55.3%
Adaptability
3.29
3.11
4.89
+48.6%
+57.2%
Role Complementarity
3.09
2.80
4.35
+40.8%
+55.4%
Subjective Experience
3.33
3.13
5.02
+50.8%
+60.4%
Overall
3.45
3.18
5.27
+52.8%
+65.7%
Appendix
Table G.4 : Category-level survey scores and relative improvements.
Category
Friedman test
PIP vs. GAMMA
PIP vs. GOAT
χ2(2)
p
p
r
p
r
Coordination
35.46
1.99×10−8
2.14×10−5
+0.62
1.42×10−7
+0.76
Predictability
39.07
3.28×10−9
6.64×10−7
+0.72
4.31×10−8
+0.82
Adaptability
33.95
4.25×10−8
1.14×10−5
+0.66
1.23×10−6
+0.72
Role Complementarity
28.31
7.12×10−7
7.10×10−5
+0.60
3.01×10−6
+0.73
Subjective Experience
32.57
8.47×10−8
1.20×10−6
+0.73
5.86×10−7
+0.76
Appendix
Table G.5 : Supporting within-subject statistical tests for post-episode survey scores. Wilcoxon p -values are uncorrected, and all PIP-vs-baseline comparisons remain significant after Holm–Bonferroni correction.
Joint attention - the human ability to share a common visual or cognitive focus with others - enables a meeting of minds that lets us coordinate even with unfamiliar partners. In this work we investigate whether equipping AI agents with a similar mechanism can enable such zero-shot coordination. We introduce Mutual Attention for zero-shot TEaming (MATE): a novel multi-agent reinforcement learning method inspired by human joint attention. MATE encourages agents to coordinate their actions by aligning their visual attention on scene-salient objects during the interaction rather than relying on arbitrary partner-dependent conventions established during training. Unlike symmetry-breaking approaches that merely prevent brittle conventions from emerging, MATE actively promotes coordination through an environment-grounded signal that is naturally shared across partners. We evaluate MATE on three benchmarks: our Card Alignment Game, designed to isolate brittle convention formation, and the more challenging Level-Based Foraging and OvercookedV2 benchmarks. Our experiments consistently show that a joint-attention-inspired signal improves coordination with unknown partners, underlining MATE's potential as a general coordination mechanism that complements and surpasses symmetry-breaking approaches.
Giulia Benintendi, Constantin Ruhdorfer, Fabian Kögel +1
Zero-shot coordination (ZSC) aims to enable agents to cooperate with independently trained partners without prior interaction, a key requirement for real-world multi-agent systems and human-AI collaboration. Existing approaches have largely emphasized increasing partner diversity during training, yet such strategies often fall short of achieving reliable generalization to unseen partners. We introduce State-Blocked Coordination (SBC), a simple yet effective framework that improves ZSC by inducing diverse interaction scenarios without direct environment modification. Specifically, SBC generates a family of virtual environments through state blocking, allowing agents to experience a wide range of suboptimal partner policies. Across multiple benchmarks, SBC demonstrates superior performance in zero-shot coordination, including strong generalization to human partners.
Mingu Kang, Sunwoo Lee, Yonghyeon Jo +1
Graduate School of Artificial Intelligence UNIST Ulsan, South Korea 44919
While AI agents are rapidly advancing from isolated tools to interactive collaborators, data-driven human-machine teaming (HMT) methods remain costly in their reliance on human interaction data across domains, teammates, and team sizes. Zero-shot coordination (ZSC) addresses this bottleneck by simulating diverse partner populations to approximate how unseen partners might behave. However, partner coverage alone is insufficient as team settings scale and communication becomes degraded. To remedy this deficiency, we propose Influence-Based Team Steering (IBTS), a framework that uses influence shaping to incentivize agents to discover diverse, high-performing team interaction patterns and further steers ongoing trajectories toward stronger learned coordination modes. We assess IBTS on Overcooked-AI in both two-agent and three-agent settings, allowing us to test whether learned coordination structure transfers beyond dyadic interaction. Our evaluation includes simulated partners, synthetic partner-style variation, and, to our knowledge, the first 30-subject Overcooked-AI HMT study involving two real human teammates and one machine teammate. Across these evaluations, IBTS improves team performance against competing baselines, highlighting the need for scaled ZSC to combine sparse-reward coordination mechanisms with partner-variation coverage rather than relying on diversity alone.