Organizations: Tianji Tec. · Shenzhen University · General Intelligence Machine · Huazhong Agricultural University · Yanbian University · South China Agricultural University · Shanghai Jiao Tong University · Hong Kong Polytechnic University · Beijing Jiaotong University · Alibaba Group
Understanding ongoing robot manipulation requires models to interpret visual observations in relation to interaction history and task progress. We introduce RoboChrono, a benchmark for streaming task understanding comprising 39 scenarios and 34,713 evaluation instances, constructed from real robot executions and complementary bare-hand human recordings. The benchmark evaluates seven tasks grouped into recognition, alignment, and temporal grounding, covering action understanding and anticipation, visual correspondence, temporal ordering, and action localization. Zero-shot evaluation of 18 vision-language models reveals substantial differences across tasks. GPT-6-Astra achieves 98.3% accuracy on Frame Matching but 68.3% on Frame Ordering, while RynnBrain1.1-122B-A10B exhibits a larger gap, reaching 95.4% and 32.9%, respectively. Input ablations on matched questions with five open-weight models further reveal distinct dependencies on visual evidence: removing visual observations reduces Current Action Recognition accuracy by 22.1 percentage points, whereas Next Action Prediction decreases by only 0.7 points. These findings show that strong visual matching does not consistently coincide with strong temporal ordering, and suggest that next-action prediction can be supported by task and action priors even when visual evidence is unavailable. RoboChrono provides a diagnostic setting for examining these differences, highlighting the need for capability-specific evaluation beyond aggregate scores when assessing task understanding in robot manipulation.
Figures & tables
Figure 1: Overview of RoboChrono. Left: Performance profiles of four representative multimodal models across the six multiple-choice evaluation tasks. Center: Distribution of 34,713 evaluation instances across seven tasks. Right: Streaming evaluation under a causal observation boundary; only observations available up to the queried moment are provided.
Benchmark / Dataset
Primary focus
Data source
Robot-mounted view
Task collection
Curation
Evaluation scale
Real robot
Head
Wrist
All teleop
Task exec.
Task QC
Human task
Manual labels
(i) Human-worn or scene-centric embodied cognition benchmarks
EgoTaskQA [ 3 ]
Human task reasoning
✗
✗
✗
✗
✓
△
✓
△
2k videos / 40k QA
VSI-Bench [ 6 ]
Spatial cognition
✗
✗
✗
✗
✗
✗
△
△
288 videos / 5k+ QA
RoboSpatial [ 12 ]
Spatial / affordance QA
✗
✗
✗
✗
✗
✗
✗
✗
1M images / 3M QA
ERQA [ 13 ]
Embodied reasoning QA
△
△
△
✗
✗
✗
✓
✓
400 QA
Table 1: Comparison of egocentric and embodied robot task benchmarks and datasets. We compare data sources, robot-mounted views, task collection and annotation protocols, and evaluation scale. RoboChrono uses real-robot executions with complementary onboard views and human-defined temporal annotations, evaluated under restricted observation histories.
Figure 2: Overview of the task coverage in RoboChrono. The benchmark includes diverse manipulation objects and action primitives drawn from 39 real-world manipulation scenarios.
Figure 3: Overview of the RoboChrono construction and evaluation pipeline. Human-annotated action intervals support task-specific question construction and unified model evaluation under a causal observation boundary.
Task
Model input
Reference target
Current action
Video prefix O≤t and answer choices
Action or state at time t
Next action
Video prefix O≤t and answer choices
Next action executed after t
Goal-conditioned next action
Video prefix O≤t , task goal, and answer choices
Next executed action under the goal
Frame matching
Query state or frame and candidate frames from O≤t
Frame matching the queried state
View matching
Query image and synchronized candidates from other views
Image from another view at the same time step
Frame ordering
Candidate frames from O≤t
Correct temporal order
Table 2: Task inputs and reference targets in RoboChrono.
Model
Size
Recognition
Alignment
Grounding
Overall
Cur. Act.
Next Act.
Next +Goal
Frame Match
Group
View Match
Frame Order
Group
Action Time
Choice Avg.
chance floor
–
0.250
0.250
0.250
0.250
0.250
0.250
0.250
0.250
–
0.250
GPT-6-Astra
–
0.894
0.764
0.836
0.983
0.873
0.675
0.683
0.679
0.721
0.805
RynnBrain1.1-122B-A10B
122B
0.752
0.511
0.588
0.954
0.708
0.536
0.329
0.442
0.512
0.628
Doubao-Seed-2.0-Lite
–
0.754
0.538
0.631
0.921
0.717
0.451
0.359
0.409
0.577
0.624
RynnBrain1.1-9B
9B
0.730
0.550
0.599
0.849
0.687
0.444
0.328
0.392
0.405
0.598
Table 3: Results on RoboChrono. Performance across seven evaluation dimensions grouped into recognition, alignment, and grounding. Group and Choice Avg. columns reproduce the aggregates from the evaluation runs over retained applicable questions. Choice Avg. excludes Action Time, which reports R@1 at tIoU ≥0.5 .
Figure 4: Diagnostic analysis of model capabilities on RoboChrono.
Evaluation dimension
GIM
Tianji
Δ
Tianji > GIM (models)
Current Action
54.1
63.7
-9.6
17 / 18
Next Action
39.9
45.0
-5.1
14 / 18
Goal-conditioned Next Action
43.9
48.1
-4.2
14 / 18
Frame Matching
57.2
66.7
-9.6
17 / 18
View Matching
32.5
35.5
-3.0
15 / 18
Frame Ordering
27.1
30.0
-2.9
17 / 18
Table 4: Task-matched comparison across robot collection platforms. We restrict GIM and Tianji gripper recordings to the 12 manipulation tasks available on both platforms and report question-weighted accuracy to isolate platform-specific effects. Scores are percentages; Δ=GIM−Tianji in percentage points. Platform accuracy is pooled over the 17 models with per-question records; the final column counts platform wins across all 18 models.
Input
Current Action
Next Action
Goal-cond. Next
Frame Match
Frame Order
View Match
Action Time
Full
62.6
40.0
49.3
46.0
29.4
34.9
30.4
Blind
40.5 (-22.1)
39.3 (-0.7)
41.1 (-8.2)
15.8 (-30.2)
27.5 (-1.9)
22.8 (-12.1)
7.6 (-22.8)
Text Only
–
–
–
27.6 (-18.4)
–
24.8 (-10.1)
–
Current Frame Only
40.0 (-22.5)
37.5 (-2.6)
46.2 (-3.0)
–
–
–
–
Shuffled Frames
53.6 (-8.9)
39.2 (-0.9)
47.6 (-1.7)
56.4 (+10.4)
–
–
12.1 (-18.3)
Identity Re-encode
62.1 (-0.4)
–
–
45.9 (+0.0)
–
–
–
Table 5: Input ablation on RoboChrono
Figure 5: Controlled analysis of collection and input effects. In (a), gains are Tianji minus GIM in percentage points, the opposite sign convention to Δ in Table 4 .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Hardware platforms used in our benchmark. We employ two real-world robotic manipulation setups with different embodiments and workspace configurations. Both platforms support VR-based human teleoperation for collecting manipulation trajectories.
Figure 7: Qualitative examples of Action Time localization. Panels (a)–(c) show the GIM examples at an enlarged scale. Ground-truth and predicted intervals are unchanged from the original examples.
Figure 8: Qualitative examples of Action Time localization (continued). Panels (d)–(f) show the Tianji examples at an enlarged scale.
Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation. These roles implicitly require models to judge not only final task success, but also how a manipulation execution is physically and temporally progressing. However, existing evaluations fail to test whether VLMs possess fine-grained process understanding. To address this gap, we present RoboProcessBench, a benchmark for process-aware understanding in vision-language robotic manipulation. RoboProcessBench decomposes such capability into two complementary dimensions, \emph{static monitoring} and \emph{dynamic reasoning}, instantiated as 12 diagnostic question families covering phase, contact, motion, coordination, primitive-local progress, temporal order, outcome, and primitive-level transitions. Built from physically grounded execution traces, the curated benchmark corpus ProcessData contains \textasciitilde 58k question-answer pairs across 260 manipulation tasks, which is further split into ProcessData-SFT and ProcessData-Eval for post-training and evaluation purposes. Extensive evaluation of various VLMs on ProcessData-Eval reveals broad limitations across 12 diagnostic task families, suggesting current models still lack robust process-aware understanding of manipulation executions. But with ProcessData-SFT, the post-trained \textit{Qwen2.5-VL-7B} and \textit{InternVL-3-8B} exhibit consistent gains on local state, motion, progress, and primitive-aware cues. These results demonstrate that RoboProcessBench serves as both an evaluation benchmark and a learnable supervision source for developing VLMs capable of monitoring and evaluating robotic manipulation processes. Project webpage: https://processbench-2026.github.io.
Dayu Xia, Yue Shi, Yao Mu +7
Shanghai AI Laboratory · Zhejiang University · Shanghai Jiao Tong University +2
A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object histories, and gestures that language leaves underspecified, yet today's manipulation benchmarks pair an instruction with a single current image, offering no way to evaluate reasoning over observed human behavior. We introduce WatchAct, a benchmark for robot manipulation grounded in observed human behavior. Each instance pairs a real-world human-action video and a language instruction with an aligned simulator scene and an executable LIBERO task, enabling scalable and reproducible evaluation. WatchAct comprises 3,000 long-horizon instances across 14 tasks in four capability domains drawn from the cognitive demands of watching another agent: parsing events (Event Grounding), recovering procedural structure (Procedural Reasoning), inferring unstated intent (Implicit Intent Inference), and tracking how the scene was changed (Episodic Reasoning). We further propose a disentangled evaluation protocol that separately measures (i)~video-to-plan reasoning by vision-language models, (ii)~policy execution under oracle plans, and (iii)~full task completion by integrated planner--policy pipelines. In both simulation and on a Franka Research 3 robot, current systems remain far from solving WatchAct. The best pipeline, Gemini-3.1-Pro with π0.5, reaches only 16.3% Success Rate (SR) in simulation and 14.0% on the real robot. Gemini-3.1-Pro attains just 36.8% Plan SR (vs. 97.1% for humans), while π0.5 reaches only 21.5% Task SR under oracle plans and drops to 10.6% on out-of-domain scenarios. Dataset and code are available at https://baiqi-li.github.io/watchact_page/.
Although the advancement of vision-language models (VLMs) has endowed robots with enhanced environmental understanding and task reasoning, a comprehensive evaluation methodology is important to advance the integration of VLMs in robotic navigation and manipulation. However, current benchmarks lack a comprehensive method to evaluate diverse robotic tasks, and evaluation metrics remain relatively constrained, making it difficult to assess the embodied capabilities of VLMs in a thorough and fine-grained manner. To address this issue, we propose RMMBench, an evaluation benchmark that requires robots to understand language instructions and perform long-horizon tasks in continuous spaces. RMMBench seamlessly integrates high- and low-level embodied tasks into a unified framework, constructing a "navigation-manipulation" task suite comprising 70 canonical task scenarios that range from localized manipulation to long-horizon composite navigation. The results reveal that leading VLMs still face major challenges in spatial localization when performing mobile manipulation tasks, and also highlight the necessity of enhancing the spatial perception capability of robots during long-horizon interactions. RMMBench can be accessed at https://mxxq-stack.github.io/rmmbench-project/
Huapeng Li, Fuxiang Feng, Jinqiu Fan +5
School of Control Science and Engineering, Shandong University · Meituan Group