RoboChrono: A Real Robot Benchmark for Streaming Task Understanding
Organizations: Tianji Tec. · Shenzhen University · General Intelligence Machine · Huazhong Agricultural University · Yanbian University · South China Agricultural University · Shanghai Jiao Tong University · Hong Kong Polytechnic University · Beijing Jiaotong University · Alibaba Group
Abstract
Understanding ongoing robot manipulation requires models to interpret visual observations in relation to interaction history and task progress. We introduce RoboChrono, a benchmark for streaming task understanding comprising 39 scenarios and 34,713 evaluation instances, constructed from real robot executions and complementary bare-hand human recordings. The benchmark evaluates seven tasks grouped into recognition, alignment, and temporal grounding, covering action understanding and anticipation, visual correspondence, temporal ordering, and action localization. Zero-shot evaluation of 18 vision-language models reveals substantial differences across tasks. GPT-6-Astra achieves 98.3% accuracy on Frame Matching but 68.3% on Frame Ordering, while RynnBrain1.1-122B-A10B exhibits a larger gap, reaching 95.4% and 32.9%, respectively. Input ablations on matched questions with five open-weight models further reveal distinct dependencies on visual evidence: removing visual observations reduces Current Action Recognition accuracy by 22.1 percentage points, whereas Next Action Prediction decreases by only 0.7 points. These findings show that strong visual matching does not consistently coincide with strong temporal ordering, and suggest that next-action prediction can be supported by task and action priors even when visual evidence is unavailable. RoboChrono provides a diagnostic setting for examining these differences, highlighting the need for capability-specific evaluation beyond aggregate scores when assessing task understanding in robot manipulation.
Figures & tables
| Benchmark / Dataset | Primary focus | Data source | Robot-mounted view | Task collection | Curation | Evaluation scale | ||||
| Real robot | Head | Wrist | All teleop | Task exec. | Task QC | Human task | Manual labels | |||
| (i) Human-worn or scene-centric embodied cognition benchmarks | ||||||||||
| EgoTaskQA [ 3 ] | Human task reasoning | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | 2k videos / 40k QA | ||
| VSI-Bench [ 6 ] | Spatial cognition | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 288 videos / 5k+ QA | ||
| RoboSpatial [ 12 ] | Spatial / affordance QA | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | 1M images / 3M QA |
| ERQA [ 13 ] | Embodied reasoning QA | ✗ | ✗ | ✗ | ✓ | ✓ | 400 QA | |||
| Task | Model input | Reference target |
| Current action | Video prefix and answer choices | Action or state at time |
| Next action | Video prefix and answer choices | Next action executed after |
| Goal-conditioned next action | Video prefix , task goal, and answer choices | Next executed action under the goal |
| Frame matching | Query state or frame and candidate frames from | Frame matching the queried state |
| View matching | Query image and synchronized candidates from other views | Image from another view at the same time step |
| Frame ordering | Candidate frames from | Correct temporal order |
| Model | Size | Recognition | Alignment | Grounding | Overall | ||||||
| Cur. Act. | Next Act. | Next +Goal | Frame Match | Group | View Match | Frame Order | Group | Action Time | Choice Avg. | ||
| chance floor | – | 0.250 | 0.250 | 0.250 | 0.250 | 0.250 | 0.250 | 0.250 | 0.250 | – | 0.250 |
| GPT-6-Astra | – | 0.894 | 0.764 | 0.836 | 0.983 | 0.873 | 0.675 | 0.683 | 0.679 | 0.721 | 0.805 |
| RynnBrain1.1-122B-A10B | 122B | 0.752 | 0.511 | 0.588 | 0.954 | 0.708 | 0.536 | 0.329 | 0.442 | 0.512 | 0.628 |
| Doubao-Seed-2.0-Lite | – | 0.754 | 0.538 | 0.631 | 0.921 | 0.717 | 0.451 | 0.359 | 0.409 | 0.577 | 0.624 |
| RynnBrain1.1-9B | 9B | 0.730 | 0.550 | 0.599 | 0.849 | 0.687 | 0.444 | 0.328 | 0.392 | 0.405 | 0.598 |
| Evaluation dimension | GIM | Tianji | Tianji GIM (models) | |
| Current Action | 54.1 | 63.7 | -9.6 | 17 / 18 |
| Next Action | 39.9 | 45.0 | -5.1 | 14 / 18 |
| Goal-conditioned Next Action | 43.9 | 48.1 | -4.2 | 14 / 18 |
| Frame Matching | 57.2 | 66.7 | -9.6 | 17 / 18 |
| View Matching | 32.5 | 35.5 | -3.0 | 15 / 18 |
| Frame Ordering | 27.1 | 30.0 | -2.9 | 17 / 18 |
| Input | Current Action | Next Action | Goal-cond. Next | Frame Match | Frame Order | View Match | Action Time |
| Full | 62.6 | 40.0 | 49.3 | 46.0 | 29.4 | 34.9 | 30.4 |
| Blind | 40.5 (-22.1) | 39.3 (-0.7) | 41.1 (-8.2) | 15.8 (-30.2) | 27.5 (-1.9) | 22.8 (-12.1) | 7.6 (-22.8) |
| Text Only | – | – | – | 27.6 (-18.4) | – | 24.8 (-10.1) | – |
| Current Frame Only | 40.0 (-22.5) | 37.5 (-2.6) | 46.2 (-3.0) | – | – | – | – |
| Shuffled Frames | 53.6 (-8.9) | 39.2 (-0.9) | 47.6 (-1.7) | 56.4 (+10.4) | – | – | 12.1 (-18.3) |
| Identity Re-encode | 62.1 (-0.4) | – | – | 45.9 (+0.0) | – | – | – |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.