A robot that assists people must often act on what a person has shown rather than said: which of several identical cartons was pointed at, or which box was handled. The plan is executed from the final scene, whereas the evidence occurs earlier, possibly on objects that have since moved. We propose instance-anchored interaction evidence (IAE), which registers every object of the final scene to its public identifier, keeps each identity through the video by backward mask propagation, and describes every frame by the geometry between hands, forearms and these instances. An evidence network trained only from task outcomes scores the instances. For pointing tasks, a grammar-constrained dynamic program trained with a structured loss decodes object-destination programs; symbolic programs handle reference disambiguation and, without learning, episodic tasks. On 1,255 WatchAct benchmark requests, scored by symbolic execution, IAE reaches 64.2% plan success on implicit-intent tasks against 27.5% for a 32B vision-language model (strict success 49.7% against 15.4%), and 46.4% against 27.0% on restoration, reversal and imitation without task-specific training. Controls with the same perception overlays, the same 32 frames, forward tracking, or a relation model trained on the same labels do not explain the gain. Given IAE's evidence as text with its meaning explained, the same language model reaches 57.4%: most of the gain comes from the instance-anchored evidence, and the explicit programs add 6.8 points at a fraction of the cost. Pointing remains the hardest case, with 16.9% strict success. The code is available at https://github.com/WeiZhou96/iae-watchact.
Figures & tables
Figure 1: Overview of instance-anchored interaction evidence (IAE). Objects of the execution scene are registered to their public identifiers and tracked backwards through the video. Hand and forearm geometry relative to each tracked instance forms the interaction evidence. An evidence network trained from task outcomes scores every candidate in every frame. For pointing tasks, a grammar-constrained dynamic program decodes object–destination operations; reference disambiguation applies a symbolic program to the learned video-level scores, and episodic tasks compare the places of tracked instances at both ends of the video without learning. Images are frames of WatchAct videos (a reference-disambiguation video, a pointing video and, for the episodic branch, the first and last frames of a restoration video); the evidence curves are schematic. Solid arrows denote inference and dashed orange arrows training-only paths.
Nonverbal cue (195)
Reference disamb. (260)
All (455)
Method
SR
Strict
SR
Strict
SR [95% CI] a
Strict [95% CI]
Δ SR b
Qwen3-VL-8B, direct
16.9
4.1
27.7
20.8
23.1 [16.3, 30.5]
13.6 [7.9, 20.0]
−4.4
Qwen3-VL-32B, direct
22.6
4.6
31.2
23.5
27.5 [20.9, 34.5]
15.4 [9.9, 21.1]
0.0
Qwen3-VL-8B + overlays
25.6
7.2
34.6
27.3
30.8 [23.1, 38.9]
18.7 [12.3, 25.3]
+3.3
Qwen3-VL-32B + overlays
21.5
3.1
34.2
24.2
28.8 [22.0, 36.0]
15.2 [9.7, 21.1]
+1.3
InternVL3.5-8B, direct
10.8
4.1
34.6
17.3
24.4 [18.5, 31.0]
11.6 [7.3, 16.5]
−3.1
Table 1: Implicit-intent tasks: plan success rate (SR) and strict success (%) on all 455 requests.
Figure 2: Examples from the test folds. Frames show the instance masks and public identifiers obtained by registration and backward tracking, the hand landmarks (white, index fingertip in yellow) and the forearm line (yellow, dashed) that IAE computes at that frame; frames are cropped at the same height for all examples. Right: frame logits of the three-seed ensemble; coloured solid curves are candidates in the goal, coloured dashed curves are candidates decoded but not in the goal (NC) or the other movable instances (RD), grey curves are the remaining candidates, and circles and squares mark decoded object and destination events (NC only). Numbers link frames to time. NC frames are the first object event and the first and last destination events; RD frames are the evidence peaks of the two cartons and the final frame. Examples were drawn at random (fixed seed) from three predefined groups: NC requests from the front camera solved strictly by IAE and not by the 32B VLM (12 requests), NC requests with the correct objects and a tray–basket confusion (18), and RD requests solved strictly by IAE and not by the 32B VLM in which the selected instance moves by more than 60 pixels (6).
Plan SR
Strict
Baseline
Δ (95% CI)
W/T/L a
p
Δ (95% CI) b
8B, direct
+41.1[31.2,50.5]
53/31/7
<10−9
+36.0[26.6,45.3]
32B, direct
+36.7[27.0,45.9]
53/27/11
<10−6
+34.3[25.7,43.1]
8B, + overlays
+33.4[24.2,42.6]
46/39/6
<10−7
+31.0[22.0,40.2]
32B, + overlays
+35.4[26.4,44.2]
55/28/8
<10−9
+34.5[25.9,42.9]
InternVL3.5-8B
+39.8[31.0,48.1]
59/27/5
<10−12
+38.0[29.0,46.8]
Table 2: Paired comparison of IAE (ensemble) with each baseline over the 91 implicit-intent activities.
NC
RD
All
Δ All h
Variant
SR
Strict
Strict
SR
Strict
SR
Strict
IAE
37.6
15.7
73.2
61.1
48.6
–
–
− arm/head rays
36.1
14.0
74.2
60.4
48.4
−0.7
−0.1
+ temporal conv. a
33.5
14.7
72.7
59.2
47.8
−1.9
−0.7
− program loss
37.3
12.3
73.6
60.7
47.3
−0.4
−1.2
Event-level task loss b
37.3
18.1
73.6
61.2
49.8
+0.1
+1.2 ∗
Table 3: Ablation and controls of IAE (three seeds, %).
Method
Imit.
Restore
Reversal
All [95% CI]
Strict
8B, direct
37.6
0.4
5.8
15.4 [10.5, 20.4]
14.1
32B, direct
37.6
16.5
26.7
27.0 [21.5, 32.6]
24.8
InternVL3.5-8B
31.7
9.8
8.9
17.5 [12.9, 22.5]
15.6
IAE
46.6
46.3
46.2
46.4 [40.2, 52.2]
46.4
Table 4: Episodic tasks without task-specific training: plan success rate (%) on 800 requests.
Figure 3: Episodic examples and instance registration. (a,b) First and last frames of two episodic requests with the tracked instance masks and public identifiers that IAE obtains at these frames; the anchor is the last frame, and arrows show the container destinations of the plan. In (a) both butter boxes are inside the baskets at the start, and each is assigned to the container nearest to its first visible position. In (b) the backward track of butter_2 , which starts inside basket_1 , coincides with the identical butter_1 in the first frame, so its place appears unchanged and its move is missed. (c) One RD activity: last frames of the three cameras with the registered identifiers; registered ground points π^(p(b)) of all instances (filled markers, one shape per camera) against their public grid positions (open squares); distance of the two identical boxes from their final positions over time in the front camera, with the frame in which butter_2 is carried. Examples were drawn at random (fixed seed) from predefined groups: restoration or reversal requests from the front camera with at least two goal pairs, solved strictly by IAE and not by the 32B VLM (48 requests); front-camera episodic requests that IAE fails and whose goal involves a container (124); RD activities with two instances of one category, successful registration on all three cameras and a moved instance (29). Frames are cropped at the same height for all cameras.
Figure 4: Reliability and consistency of IAE on the implicit-intent tasks. (a,b) Success of the committed requests when requests are ranked by the confidence of the three-seed ensemble (NC: program margin μ ; RD: selection margin minc∣sc∣ ); shaded bands are 95% activity-bootstrap intervals of strict success, and curves start at ten committed requests. The strip under each panel marks every request in ranked order, filled when strictly successful. Dotted and dash-dotted grey curves are the oracle ranking (all strictly successful requests first) and a random ranking; the areas under the three strict curves are given in each panel. Dashed coloured lines give the strict success of two direct baselines, which provide no confidence. (c) Per-activity difference in plan SR between IAE and Qwen3-VL-32B direct planning; black bars give the mean and its 95% bootstrap interval.
Camera
Reference
Method
Front
Side
Obl.
Camera
Person
Nonverbal cue
32B, direct
28.2
14.1
28.2
25.6
20.5
8B, direct + overlays
26.9
23.1
28.2
25.6
25.6
IAE (ensemble)
46.2
41.0
38.5
43.6
41.9
Reference disambiguation
Table 5: Plan success rate (%) by camera and spatial reference.
Figure 5: Where plans fail on the implicit-intent tasks. (a,b) Outcome of every goal pair: the object is moved to its goal destination (correct pair), moved elsewhere (wrong destination) or not moved (missed). Circles give the share of requests whose plan additionally moves an object that is not in the goal, which the official metric does not penalize but strict success does. (c) Types of wrong destinations on NC for IAE (ensemble, filled) and Qwen3-VL-32B direct planning (open); numbers are counts for IAE/ 32B.
Figure 6: Data efficiency and sensitivity of IAE on the implicit-intent tasks (plan SR, solid; strict success, dashed; all 455 requests in orange, NC in blue). (a–c) Retraining with a fraction of the training activities of each fold, a different number k of frames pooled into the video score, and a different weight γ of the program-level loss; lines are means over three seeds, bands one standard deviation, and small dots the individual seeds. (d) Test-time decoding cost λ applied to the frame logits of the three-seed ensemble without retraining; RD does not depend on λ . Dotted lines mark the settings used elsewhere in the paper.
NC
RD
All
Episodic
Layout at test time
IDs changed a
SR
SR
SR
Strict
SR
Exact layout
0.0
37.6
78.7
61.1
48.6
46.4
Noise σ=0.25 cell
4.5
37.1
78.5
60.7
48.2
44.0
Noise σ=0.5 cell
15.3
35.2
72.2
56.3
44.5
38.0
Noise σ=1 cell
34.2
34.0
60.1
48.9
37.1
30.7
Random identities b
51.4
32.8
49.2
42.2
30.5
27.1
Table 6: Dependence on the public layout used for registration (%, three seeds).
Stage
Model
Time (s)
Anchor-frame detection
Grounding DINO-T
1.1
0.3
Layout registration
least squares
< 0.01
< 0.01
Mask propagation
SAM 2.1-T
20.2
19.0
Per-frame person detection
Grounding DINO-T
36.1
35.3
Hand keypoints
MediaPipe Hands
8.0
10.2
Body keypoints
RTMPose-m
4.4
4.5
Table 7: Computation for two test videos of 34.5 s and 33.3 s (one RTX 5880 Ada GPU; keypoints on CPU).
A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object histories, and gestures that language leaves underspecified, yet today's manipulation benchmarks pair an instruction with a single current image, offering no way to evaluate reasoning over observed human behavior. We introduce WatchAct, a benchmark for robot manipulation grounded in observed human behavior. Each instance pairs a real-world human-action video and a language instruction with an aligned simulator scene and an executable LIBERO task, enabling scalable and reproducible evaluation. WatchAct comprises 3,000 long-horizon instances across 14 tasks in four capability domains drawn from the cognitive demands of watching another agent: parsing events (Event Grounding), recovering procedural structure (Procedural Reasoning), inferring unstated intent (Implicit Intent Inference), and tracking how the scene was changed (Episodic Reasoning). We further propose a disentangled evaluation protocol that separately measures (i)~video-to-plan reasoning by vision-language models, (ii)~policy execution under oracle plans, and (iii)~full task completion by integrated planner--policy pipelines. In both simulation and on a Franka Research 3 robot, current systems remain far from solving WatchAct. The best pipeline, Gemini-3.1-Pro with π0.5, reaches only 16.3% Success Rate (SR) in simulation and 14.0% on the real robot. Gemini-3.1-Pro attains just 36.8% Plan SR (vs. 97.1% for humans), while π0.5 reaches only 21.5% Task SR under oracle plans and drops to 10.6% on out-of-domain scenarios. Dataset and code are available at https://baiqi-li.github.io/watchact_page/.
Humans can learn a new manipulation task from one or two demonstrations and then perform it in a new room, with new objects, under new constraints. Modern robot imitation learning, in contrast, typically needs hundreds to thousands of demonstrations and still degrades under modest shifts in layout, geometry, object set or task constraints. We argue this gap is not just about data, but also about the level of abstraction at which learning occurs; generalization requires inferring the latent intent underlying why a demonstrator behaved in a certain way, rather than reproducing how they moved. We present Rational Inverse Reasoning (RIR), which casts few-shot imitation as inference over latent explanation programs: compact, executable descriptions of intent that map an object-centric scene to a structured task-and-motion-planning (TAMP) specification of goals, subgoals and constraints. A vision-language model proposes candidate programs, and a hierarchical planner supplies a bounded-rational likelihood. By combining VLM program proposals, and planner-grounded feedback, RIR iteratively refines the candidate set to approximate a posterior over concise, executable programs. On a 2D reasoning benchmark and a real Franka FR3, RIR recovers transferable task structure from as little as one demonstration. Generalizing to substantially new layouts and object sets, RIR outperforms VLM-planning baselines that lack explicit rationality and planning-grounded inference, increasing downstream success rate by 34 and 28 percentage points in the one- and three-shot settings.
Ben Zandonati, Tomás Lozano-Pérez, Leslie Pack Kaelbling
Understanding natural human instructions is crucial for deploying robots in human-centric environments. We study multimodal instruction grounding, where language and gesture provide complementary but uncertain cues. We present MIGU, a modular framework that combines semantic and geometric evidence into a unified grounding belief and connects it to manipulation planning. MIGU constructs a 3D geometric likelihood by propagating viewing-direction and depth uncertainty through eye-finger geometry while accounting for hand-direction estimation error. A vision-language model (VLM) provides semantic priors over candidate objects and regions, which are combined with the geometric likelihood through Bayes-inspired fusion. The resulting belief supports behavior planning to either proceed directly to downstream planning or request clarification. Grounded targets then define goals for mobile manipulation and tabletop task-and-motion planning. On a real-world benchmark, MIGU outperforms all evaluated baselines, while ablations support the benefit of explicit multimodal uncertainty modeling. Project website: multimodal-instruction.github.io
Mingke Lu, Anxing Xiao, David Hsu
School of Computing and Smart Systems Institute, National University of Singapore, Singapore · Department of Electrical and Computer Engineering, University of California, Los Angeles, Los Angeles, CA, USA