cs.ROOct 8, 2026

Instance-anchored interaction evidence: Grounding robot plans in human pointing and handling

Authors: Xinliang Xiao, Bowen Yang, Wenjing Zhang, Li Yang, Wei Zhou

Organizations: Nanjing University of Science and Technology, Nanjing, 210094, China

Abstract

A robot that assists people must often act on what a person has shown rather than said: which of several identical cartons was pointed at, or which box was handled. The plan is executed from the final scene, whereas the evidence occurs earlier, possibly on objects that have since moved. We propose instance-anchored interaction evidence (IAE), which registers every object of the final scene to its public identifier, keeps each identity through the video by backward mask propagation, and describes every frame by the geometry between hands, forearms and these instances. An evidence network trained only from task outcomes scores the instances. For pointing tasks, a grammar-constrained dynamic program trained with a structured loss decodes object-destination programs; symbolic programs handle reference disambiguation and, without learning, episodic tasks. On 1,255 WatchAct benchmark requests, scored by symbolic execution, IAE reaches 64.2% plan success on implicit-intent tasks against 27.5% for a 32B vision-language model (strict success 49.7% against 15.4%), and 46.4% against 27.0% on restoration, reversal and imitation without task-specific training. Controls with the same perception overlays, the same 32 frames, forward tracking, or a relation model trained on the same labels do not explain the gain. Given IAE's evidence as text with its meaning explained, the same language model reaches 57.4%: most of the gain comes from the instance-anchored evidence, and the explicit programs add 6.8 points at a fraction of the cost. Pointing remains the hardest case, with 16.9% strict success. The code is available at https://github.com/WeiZhou96/iae-watchact.

Figures & tables

Explore similar work

Jun 24, 2026cs.RO

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation

A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object histories, and gestures that language leaves underspecified, yet today's manipulation benchmarks pair an instruction with a single current image, offering no way to evaluate reasoning over observed human behavior. We introduce WatchAct, a benchmark for robot manipulation grounded in observed human behavior. Each instance pairs a real-world human-action video and a language instruction with an aligned simulator scene and an executable LIBERO task, enabling scalable and reproducible evaluation. WatchAct comprises 3,000 long-horizon instances across 14 tasks in four capability domains drawn from the cognitive demands of watching another agent: parsing events (Event Grounding), recovering procedural structure (Procedural Reasoning), inferring unstated intent (Implicit Intent Inference), and tracking how the scene was changed (Episodic Reasoning). We further propose a disentangled evaluation protocol that separately measures (i)~video-to-plan reasoning by vision-language models, (ii)~policy execution under oracle plans, and (iii)~full task completion by integrated planner--policy pipelines. In both simulation and on a Franka Research 3 robot, current systems remain far from solving WatchAct. The best pipeline, Gemini-3.1-Pro with π0.5π_{0.5}, reaches only 16.3% Success Rate (SR) in simulation and 14.0% on the real robot. Gemini-3.1-Pro attains just 36.8% Plan SR (vs. 97.1% for humans), while π0.5π_{0.5} reaches only 21.5% Task SR under oracle plans and drops to 10.6% on out-of-domain scenarios. Dataset and code are available at https://baiqi-li.github.io/watchact_page/.
Aug 12, 2025cs.RO

Rational Inverse Reasoning: Few-Shot Imitation by Inferring Intent through Planning

Humans can learn a new manipulation task from one or two demonstrations and then perform it in a new room, with new objects, under new constraints. Modern robot imitation learning, in contrast, typically needs hundreds to thousands of demonstrations and still degrades under modest shifts in layout, geometry, object set or task constraints. We argue this gap is not just about data, but also about the level of abstraction at which learning occurs; generalization requires inferring the latent intent underlying why a demonstrator behaved in a certain way, rather than reproducing how they moved. We present Rational Inverse Reasoning (RIR), which casts few-shot imitation as inference over latent explanation programs: compact, executable descriptions of intent that map an object-centric scene to a structured task-and-motion-planning (TAMP) specification of goals, subgoals and constraints. A vision-language model proposes candidate programs, and a hierarchical planner supplies a bounded-rational likelihood. By combining VLM program proposals, and planner-grounded feedback, RIR iteratively refines the candidate set to approximate a posterior over concise, executable programs. On a 2D reasoning benchmark and a real Franka FR3, RIR recovers transferable task structure from as little as one demonstration. Generalizing to substantially new layouts and object sets, RIR outperforms VLM-planning baselines that lack explicit rationality and planning-grounded inference, increasing downstream success rate by 3434 and 2828 percentage points in the one- and three-shot settings.
Sep 21, 2026cs.RO

MIGU: Multimodal Instruction Grounding under Uncertainty for Manipulation Planning

Understanding natural human instructions is crucial for deploying robots in human-centric environments. We study multimodal instruction grounding, where language and gesture provide complementary but uncertain cues. We present MIGU, a modular framework that combines semantic and geometric evidence into a unified grounding belief and connects it to manipulation planning. MIGU constructs a 3D geometric likelihood by propagating viewing-direction and depth uncertainty through eye-finger geometry while accounting for hand-direction estimation error. A vision-language model (VLM) provides semantic priors over candidate objects and regions, which are combined with the geometric likelihood through Bayes-inspired fusion. The resulting belief supports behavior planning to either proceed directly to downstream planning or request clarification. Grounded targets then define goals for mobile manipulation and tabletop task-and-motion planning. On a real-world benchmark, MIGU outperforms all evaluated baselines, while ablations support the benefit of explicit multimodal uncertainty modeling. Project website: multimodal-instruction.github.io