Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that links actions on affordance parts to semantic and geometric state changes. By representing observed and anticipated transitions in the same form, it provides a shared basis for understanding and planning. We construct ChronoGraphBench through an automatic data engine that converts human-interaction videos and simulated robot trajectories into graph-annotated questions for training and evaluating Vision-Language Models (VLMs) on both tasks. Using these annotations, we train ChronoGraphVLM by adapting pretrained VLMs in two stages. Graph-as-Chain-of-Thought supervised fine-tuning teaches the models to reconstruct observed transitions and predict future ones as graph traces before answering. Subsequent joint 4D graph reinforcement learning directly rewards graph properties and answer correctness. Experiments across model scales show improvements over the corresponding pretrained baselines and zero-shot transfer to VLM4D. Real-world demonstrations further show that graph-based planning and affordance grounding support mobile manipulation through existing robot skills without additional fine-tuning.
Figures & tables
Figure 1: ChronoGraph connects 4D interaction understanding with spatially grounded planning. ChronoGraphVLM generates functional 4D scene graphs linking affordance-level actions to semantic and geometric state changes for interaction-grounded VQA and planning. Its predicted plans and affordance grounding guide zero-shot real-world mobile manipulation.
Figure 2: ChronoGraphBench annotation pipeline. We convert interaction recordings into functional 4D scene graphs linking affordance-level actions to semantic and geometric state changes. The pipeline combines semantic graph initialization, object and affordance grounding, human verification, and graph refinement.
4D Interaction Understanding
Spatially Grounded Planning
Model
Action Recognition
Affordance Recognition
Spatial Reasoning
State Change Recognition
Avg.
Action Prediction
Affordance Prediction
Spatial Prediction
State Change Prediction
2D/3D Grounding
Avg.
Overall Avg.
Commercial Models
GPT 6 Astra
93.8
72.6
75.5
88.1
85.5
85.4/85.4
81.1/81.7
78.1/69.2
85.3/96.6
—/67.9
81.3/74.2
79.4
GPT 5.6 Sol
86.8
72.6
69.6
83.5
80.1
78.0/85.4
78.4/78.5
75.0/57.8
88.2/93.1
—/61.2
78.4/68.1
73.9
GPT 5.6 Terra
84.9
74.2
64.1
77.8
76.5
75.6/82.6
67.6/83.9
65.6/54.0
61.8/93.1
—/61.5
67.3/67.4
70.9
Gemini 3.1 Pro
85.3
74.2
65.2
82.5
78.2
75.6/86.1
73.0/84.9
67.7/63.5
88.2/100.0
—/61.0
73.6/70.5
73.8
Table 1: Main comparison on ChronoGraphBench for 4D interaction understanding and spatially grounded planning. Planning scores are reported as history-conditioned/static-conditioned , with 2D/3D grounding evaluated only in the static-conditioned setting. The overall average is the unweighted mean of per-question scores across all settings.
Method
Acc. (%) ↑
Qwen3.5-9B
60.5
Qwen3.5-4B
53.7
Ours-9B
67.1
Ours-4B
61.0
Table 2: Zero-shot transfer to VLM4D. We report multiple-choice accuracy on the egocentric split.
4D Interaction Understanding
Spatially Grounded Planning
VLM4D
Method
Action Recognition
Affordance Recognition
Spatial Reasoning
State Change Recognition
Avg.
Action Prediction
Affordance Prediction
Spatial Prediction
State Change Prediction
2D/3D Grounding
Avg.
Overall Avg.
Acc. (%)
(a) Graph-Guided Reasoning
Qwen3.5-9B (Direct)
62.8
59.7
38.0
60.3
55.3
48.8/71.5
45.9/79.6
55.2/41.7
50.0/89.7
—/29.1
51.4/47.9
51.2
60.5
Qwen3.5-9B + Naive Thinking
68.3
65.6
59.0
59.7
63.8
46.2/76.9
44.4/77.2
45.5/47.6
76.4/92.7
—/26.0
49.5/50.6
55.8
62.2
Qwen3.5-9B + Graph-as-CoT
67.8
67.7
50.0
70.1
63.8
51.2/80.6
45.9/77.4
58.3/49.8
50.0/87.9
—/42.4
53.3/56.7
59.0
63.9
(b) Supervised Fine-Tuning
Table 3: Ablations of graph-guided reasoning and training with Qwen3.5-9B. ChronoGraphBench planning scores are reported as history-conditioned / static-conditioned . The final column reports zero-shot accuracy on the egocentric split of VLM4D.
Figure 3: Zero-shot real-world mobile manipulation with a Boston Dynamics Spot robot. ChronoGraphVLM predicts sub-actions and updates affordance grounding before each interaction, guiding tasks of increasing horizon through existing robot skills.
University of Arkansas, Fayetteville, AR, USA. · Max Planck Research School for Intelligent Systems and the University of Stuttgart, Stuttgart, Germany. · Center of AI Research, VinUniversity, VietNam. +2
Nanjing University, Nanjing, China · Institute of Automation, Chinese Academy of Sciences, Beijing, China · University of Chinese Academy of Sciences, Beijing, China