Robots operating in the physical world will increasingly need to coordinate with other robots, particularly in manipulation tasks where an object may be too large or heavy for a single robot to carry alone. Physical limitations caused by hardware degradation or actuator faults can restrict the actions a robot can reliably execute, yet these limitations may be unknown to its partner. We study whether a helper can infer a robot partner's physical constraints from observing it coordinate with another robot, then use the inferred capability to coordinate with the same partner on a new task. This is difficult because a demonstration shows what the constrained robot did, but not what it could have done. In physically coupled tasks, the other robot may also compensate for its limitations, making those limitations difficult to identify from the constrained robot's behavior alone. Our key insight is that these constraints shape the joint behavior of the team, making the actions of both robots informative about the constrained partner's capability. We introduce Watch, Infer, Coordinate, a benchmark spanning three physically coupled manipulation settings, together with an inference approach that scores candidate constraints using observed joint behavior. Across all three settings, our method substantially improves constraint inference and zero-shot coordination, approaching an oracle with access to the true constraints.
Figures & tables
Figure 1: Overview of our problem setting. A helper must coordinate with a constrained partner whose physical capabilities are initially unknown. Rather than interacting with the partner directly, the helper first watches it coordinate with another robot and uses this demonstration to infer constraints on its low-level actions, such as joint position, joint velocity, or base-direction limits. The helper then replaces the demonstrator and must coordinate with the same constrained partner on a new task. Although the required coordination changes between tasks, the partner’s physical constraints remain the same, allowing the helper to transfer what it learned about the partner’s capabilities to the new task.
Work
Prior-Trajectory Inference
Multi-Agent Trajectories
Low-Level Partner Constraints
Continuous Control
Physically Coupled Manipulation
Transfer to New Task
CHORUS ( Doshi et al., 2026 )
✗
✗
✗
✓
✓
✗
Watch-and-Help ( Puig et al., 2021 )
✓
✗
✗
✗
✗
✗
TICC-POMDP ( Lee et al., 2020 )
✗
✗
✗
✗
✗
✓
Smart Help ( Cao et al., 2024 )
✗
✗
✗
✗
✗
✗
CHAIC ( Du et al., 2024 )
✗
✗
✗
✗
✓
✗
CE-CM-Div ( Tisnikar et al., 2026 )
✓
✓
✗
✗
✗
✓
Table 1: Comparison with related work. Prior-Trajectory Inference indicates that episode-specific information is inferred from trajectories observed before the downstream task. Multi-Agent Trajectories indicates that the trajectories used for inference contain behavior from multiple interacting agents. Low-Level Partner Constraints indicates that the inferred quantity specifies low-level constraints such as limits on the partner’s joint positions, joint velocities, or allowable motion directions, rather than task-level capabilities such as reachable height or liftable weight. Continuous Control indicates continuous-valued physical control rather than symbolic or macro actions. Physically Coupled Manipulation indicates that multiple agents jointly manipulate the same object such that their actions mechanically combine to determine its motion. Transfer to New Task indicates that the inferred information is subsequently used on a distinct task or goal requiring different behavior.
Figure 2: Overview of our physically coupled partner-adaptation benchmark. (A) During the observation stage, the helper watches a demonstrator coordinate with the constrained agent on an obstacle-free task. The labels show physical constraints on the action space of the constrained agent, which are dependent on the setting. (B) At test time, the helper replaces the demonstrator and coordinates zero-shot with the same constrained agent on a new task with obstacles. The benchmark includes three settings of increasing physical complexity: 2D rod carrying, fixed-base dual-UR5 carrying, and mobile dual-UR5 carrying. Across all settings, the constrained agent’s physical constraints remain fixed while the task and helper change.
2D Rod
Fixed-Base Dual-UR5
Mobile Dual-UR5
N
CE-CM-Div
Ours, partner only
Ours
CE-CM-Div
Ours, partner only
Ours
CE-CM-Div
Ours, partner only
Ours
1
0.304±0.038
0.104±0.015
0.058 ± 0.017
0.074±0.000
0.068±0.006
0.031 ± 0.012
0.306±0.028
0.201±0.045
0.190 ± 0.060
2
0.340±0.041
0.031±0.009
0.003 ± 0.002
0.074±0.000
0.037±0.000
0.019 ± 0.000
0.278±0.028
0.148±0.042
0.130 ± 0.060
4
0.348±0.031
0.002 ± 0.002
0.002 ± 0.002
0.074±0.000
0.037±0.000
0.019 ± 0.000
0.250±0.000
0.090±0.006
0.069 ± 0.005
8
0.302±0.011
0.000 ± 0.000
0.002±0.002
0.074±0.000
0.043±0.006
0.019 ± 0.011
0.250±0.000
0.042±0.000
0.009 ± 0.000
Table 2: Constraint-inference accuracy at N∈{1,2,4,8} demonstrations. We report normalized Hamming distance between the inferred and true constraints. “Ours, partner only” uses only the constrained agent’s behavior for inference and excludes the demonstrator’s behavior. Values are mean ± standard error; lower is better. Bold indicates the best result within each robot setting and row, including ties.
2D Rod
Fixed-Base Dual-UR5
Mobile Dual-UR5
Task
Method
Success ↑
SNA ↑
Success ↑
SNA ↑
Success ↑
SNA ↑
Obstacle-free control
Capacity-blind + CEM
100.0 ± 0.0
91.4±0.9
0.0±0.0
0.0±0.0
66.7±0.0
55.2±0.0
CE-CM-Div + CEM
100.0 ± 0.0
90.4±1.6
55.6±11.1
30.5±8.3
66.7±0.0
54.3±0.0
Behavioral cloning
66.7±0.0
52.5±6.1
22.2±11.1
7.5±3.8
66.7±0.0
31.1±0.0
Oracle + CEM
100.0±0.0
91.4±0.2
100.0±0.0
44.8±1.8
100.0±0.0
80.2±0.0
Ours + CEM
100.0 ± 0.0
92.4 ± 0.4
100.0 ± 0.0
43.9 ± 0.8
100.0 ± 0.0
80.2 ± 0.0
Table 3: Planning performance on the obstacle-free control and new tasks with obstacles. Our method, CE-CM-Div, and behavioral cloning use N=8 demonstrations. Capacity-blind CEM uses no demonstrations, while Oracle CEM receives the true constraints. Success and SNA are percentages reported as mean ± standard error over three seeds; higher is better. Bold indicates the best non-oracle result for each task and metric, including ties.
Figure 3: Qualitative examples of zero-shot coordination using our method. (A) Representative successful trajectories for 2D rod carrying, fixed-base dual-UR5 carrying, and mobile dual-UR5 carrying, shown from left to right over time. (B) Example of how an inferred physical constraint changes test-time planning. Ignoring the constrained agent’s limitation leads to a pose it cannot reach, while incorporating the inferred constraint causes the planner to select a feasible alternative.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Setting
Horizon
Population
Elites
Iterations
2D Rod
32
64
8
3
Fixed-Base Dual-UR5
4
16
2
2
Mobile Dual-UR5
4
4
1
1
Appendix
Table 4: CEM-MPC hyperparameters used for zero-shot coordination.
Humans combine reasoning and motor control to solve complex manipulation tasks under diverse constraints. They build an understanding of the physical world that helps them convert reasoning into actions and quickly adapt to new scenes, tasks, and rules. We refer to this capability as intuitive manipulation. Existing benchmarks fail to capture this integration: they evaluate physical reasoning in isolation from execution, or measure policy performance without requiring explicit reasoning. We introduce IMBENCH, a benchmark designed to evaluate intuitive manipulation as an integrated capability spanning perception, physical reasoning, action generation, and iterative execution. Our tasks require models to infer task-relevant physical structure and generate feasible action sequences under explicit constraints, including contact-rich manipulation, tool use, and multi-stage dependencies. We introduce a benchmark of 35 tasks, 14K filtered trajectories, and scalable tools for generating diverse scenarios. Experiments reveal a consistent gap: vision language models show partial physical reasoning ability but fail to produce executable plans, while state-of-the-art vision-language-action models struggle to satisfy task constraints and generalize across scenarios. These results identify intuitive manipulation as a missing axis in current foundation models and generalist robot policies, and position IMBENCH as a step toward evaluating and enabling more integrated, adaptive physical intelligence.
A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object histories, and gestures that language leaves underspecified, yet today's manipulation benchmarks pair an instruction with a single current image, offering no way to evaluate reasoning over observed human behavior. We introduce WatchAct, a benchmark for robot manipulation grounded in observed human behavior. Each instance pairs a real-world human-action video and a language instruction with an aligned simulator scene and an executable LIBERO task, enabling scalable and reproducible evaluation. WatchAct comprises 3,000 long-horizon instances across 14 tasks in four capability domains drawn from the cognitive demands of watching another agent: parsing events (Event Grounding), recovering procedural structure (Procedural Reasoning), inferring unstated intent (Implicit Intent Inference), and tracking how the scene was changed (Episodic Reasoning). We further propose a disentangled evaluation protocol that separately measures (i)~video-to-plan reasoning by vision-language models, (ii)~policy execution under oracle plans, and (iii)~full task completion by integrated planner--policy pipelines. In both simulation and on a Franka Research 3 robot, current systems remain far from solving WatchAct. The best pipeline, Gemini-3.1-Pro with π0.5, reaches only 16.3% Success Rate (SR) in simulation and 14.0% on the real robot. Gemini-3.1-Pro attains just 36.8% Plan SR (vs. 97.1% for humans), while π0.5 reaches only 21.5% Task SR under oracle plans and drops to 10.6% on out-of-domain scenarios. Dataset and code are available at https://baiqi-li.github.io/watchact_page/.
We present a structured large-language-model (LLM) architecture for zero-shot human--robot coordination in a cooperative construction task with private goal views. Guided by a Dec-POMDP formulation, the architecture decomposes decision-making into (i) action-conditioned Theory-of-Mind (ToM) inference, (ii) hierarchical planning, (iii) conversation interpretation, (iv) action verification, and (v) feedback-based replanning. We compare the proposed method with an ablation without ToM inference and a multi-agent reinforcement-learning policy trained offline over many goal pairs. In human-participant experiments, the proposed method required fewer interaction steps and yielded higher post-interaction trust ratings than both baselines. These results suggest that systematically decomposing the team decision problem, using LLMs as tractable surrogates for otherwise intractable inference and planning computations, and retaining conventional verification for physical feasibility can improve both task coordination and the human experience.
Dong Hae Mangalindan, Anand Gokhale, Francesco Bullo +1
Department of Electrical and Computer Engineering, Michigan State University, East Lansing, MI 48824 · Center for Control, Dynamical Systems, and Computation, UC Santa Barbara, Santa Barbara, CA 93106 USA