Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that links actions on affordance parts to semantic and geometric state changes. By representing observed and anticipated transitions in the same form, it provides a shared basis for understanding and planning. We construct ChronoGraphBench through an automatic data engine that converts human-interaction videos and simulated robot trajectories into graph-annotated questions for training and evaluating Vision-Language Models (VLMs) on both tasks. Using these annotations, we train ChronoGraphVLM by adapting pretrained VLMs in two stages. Graph-as-Chain-of-Thought supervised fine-tuning teaches the models to reconstruct observed transitions and predict future ones as graph traces before answering. Subsequent joint 4D graph reinforcement learning directly rewards graph properties and answer correctness. Experiments across model scales show improvements over the corresponding pretrained baselines and zero-shot transfer to VLM4D. Real-world demonstrations further show that graph-based planning and affordance grounding support mobile manipulation through existing robot skills without additional fine-tuning.
Figures & tables
Figure 1: ChronoGraph connects 4D interaction understanding with spatially grounded planning. ChronoGraphVLM generates functional 4D scene graphs linking affordance-level actions to semantic and geometric state changes for interaction-grounded VQA and planning. Its predicted plans and affordance grounding guide zero-shot real-world mobile manipulation.
Figure 2: ChronoGraphBench annotation pipeline. We convert interaction recordings into functional 4D scene graphs linking affordance-level actions to semantic and geometric state changes. The pipeline combines semantic graph initialization, object and affordance grounding, human verification, and graph refinement.
4D Interaction Understanding
Spatially Grounded Planning
Model
Action Recognition
Affordance Recognition
Spatial Reasoning
State Change Recognition
Avg.
Action Prediction
Affordance Prediction
Spatial Prediction
State Change Prediction
2D/3D Grounding
Avg.
Overall Avg.
Commercial Models
GPT 6 Astra
93.8
72.6
75.5
88.1
85.5
85.4/85.4
81.1/81.7
78.1/69.2
85.3/96.6
—/67.9
81.3/74.2
79.4
GPT 5.6 Sol
86.8
72.6
69.6
83.5
80.1
78.0/85.4
78.4/78.5
75.0/57.8
88.2/93.1
—/61.2
78.4/68.1
73.9
GPT 5.6 Terra
84.9
74.2
64.1
77.8
76.5
75.6/82.6
67.6/83.9
65.6/54.0
61.8/93.1
—/61.5
67.3/67.4
70.9
Gemini 3.1 Pro
85.3
74.2
65.2
82.5
78.2
75.6/86.1
73.0/84.9
67.7/63.5
88.2/100.0
—/61.0
73.6/70.5
73.8
Table 1: Main comparison on ChronoGraphBench for 4D interaction understanding and spatially grounded planning. Planning scores are reported as history-conditioned/static-conditioned , with 2D/3D grounding evaluated only in the static-conditioned setting. The overall average is the unweighted mean of per-question scores across all settings.
Method
Acc. (%) ↑
Qwen3.5-9B
60.5
Qwen3.5-4B
53.7
Ours-9B
67.1
Ours-4B
61.0
Table 2: Zero-shot transfer to VLM4D. We report multiple-choice accuracy on the egocentric split.
4D Interaction Understanding
Spatially Grounded Planning
VLM4D
Method
Action Recognition
Affordance Recognition
Spatial Reasoning
State Change Recognition
Avg.
Action Prediction
Affordance Prediction
Spatial Prediction
State Change Prediction
2D/3D Grounding
Avg.
Overall Avg.
Acc. (%)
(a) Graph-Guided Reasoning
Qwen3.5-9B (Direct)
62.8
59.7
38.0
60.3
55.3
48.8/71.5
45.9/79.6
55.2/41.7
50.0/89.7
—/29.1
51.4/47.9
51.2
60.5
Qwen3.5-9B + Naive Thinking
68.3
65.6
59.0
59.7
63.8
46.2/76.9
44.4/77.2
45.5/47.6
76.4/92.7
—/26.0
49.5/50.6
55.8
62.2
Qwen3.5-9B + Graph-as-CoT
67.8
67.7
50.0
70.1
63.8
51.2/80.6
45.9/77.4
58.3/49.8
50.0/87.9
—/42.4
53.3/56.7
59.0
63.9
(b) Supervised Fine-Tuning
Table 3: Ablations of graph-guided reasoning and training with Qwen3.5-9B. ChronoGraphBench planning scores are reported as history-conditioned / static-conditioned . The final column reports zero-shot accuracy on the egocentric split of VLM4D.
Figure 3: Zero-shot real-world mobile manipulation with a Boston Dynamics Spot robot. ChronoGraphVLM predicts sub-actions and updates affordance grounding before each interaction, guiding tasks of increasing horizon through existing robot skills.
Vision-Language-Action (VLA) models promise generalist robot manipulation, but are typically trained and deployed as short-horizon policies that assume the latest observation is sufficient for action reasoning. This assumption breaks in non-Markovian long-horizon tasks, where task-relevant evidence can be occluded or appear only earlier in the trajectory, and where clutter and distractors make fine-grained visual grounding brittle. We present CodeGraphVLP, a hierarchical framework that enables reliable long-horizon manipulation by combining a persistent semantic-graph state with an executable code-based planner and progress-guided visual-language prompting. The semantic-graph maintains task-relevant entities and relations under partial observability. The synthesized planner executes over this semantic-graph to perform efficient progress checks and outputs a subtask instruction together with subtask-relevant objects. We use these outputs to construct clutter-suppressed observations that focus the VLA executor on critical evidence. On real-world non-Markovian tasks, CodeGraphVLP improves task completion over strong VLA baselines and history-enabled variants while substantially lowering planning latency compared to VLM-in-the-loop planning. We also conduct extensive ablation studies to confirm the contributions of each component.
Khoa Vo, Sieu Tran, Taisei Hanyu +8
University of Arkansas, Fayetteville, AR, USA. · Max Planck Research School for Intelligent Systems and the University of Stuttgart, Stuttgart, Germany. · Center of AI Research, VinUniversity, VietNam. +2
Long-horizon robotic manipulation requires vision-language-action (VLA) models to track scene states and their evolution beyond the current observation. However, simply conditioning policies on observation history does not guarantee that the history is effectively utilized: action supervision constrains what the policy should do, but only indirectly constrains what its history representations should retain. To resolve this, we present Temporal Forcing, a 4D representation alignment framework that explicitly supervises latent temporal states and their transitions. Specifically, we first introduce a history pathway that compresses past observations into compact latent tokens. We then align these tokens and current-frame features with geometric targets from a pretrained 4D foundation model, providing direct supervision at both the state and transition levels. The 4D foundation model and alignment heads are used only for training-time supervision. Temporal Forcing improves average success from 96.6% to 98.8% on LIBERO, with the largest gain on LIBERO-Long (93.8% to 97.2%), and from 53.5% to 62.8% across twelve RoboTwin 2.0 tasks. Furthermore, Temporal Forcing increases full-task success from 20.0% to 43.3% on a physical multi-stage hidden-placement task. Controlled experiments show that 4D representation alignment is crucial for making observation history beneficial to the model. Code will be publicly available.
Xingyu Ding, Yuzhong Zhao, Chunhai Zhao +3
Nanjing University, Nanjing, China · Institute of Automation, Chinese Academy of Sciences, Beijing, China · University of Chinese Academy of Sciences, Beijing, China
Vision Language Models (VLMs) have shown promising planning capabilities, yet their success remains confined to the text domain, leaving visual decision-making relatively underexplored. Addressing this gap, we introduce Corrective Sequence Planning (CoSPlan) benchmark, where VLMs must plan a sequence of visual actions from an initial scene to a target scene. CoSPlan evaluates models on their ability to imagine and execute a coherent set of visual steps required to reach the goal (Step Completion). To prevent any shortcuts that simply describe the final scene, we introduce an erroneous action in decision-making, which must be detected (Error Detection) and corrected to reach the goal, enabling a deeper understanding of the task. CoSPlan spans across 4 tasks: maze navigation, block re-arrangement, image reconstruction, and object re-organization. Despite using advanced reasoning strategies such as Chain-of-Thought and Scene Graphs, VLMs struggle on CoSPlan, while still showing promising performance in the text domain. Addressing this, we propose Scene Graph Incremental updates (SGI), a novel training-free method to transform images into `textual' scene graphs, enabling step-by-step reasoning through iterative scene graph refinement. SGI yields an average of ~4.4% improvement on CoSPlan w/ generalization on PlanBench and VQA. Link for solving puzzles on the project page.
Shresth Grover, Priyank Pathak, Akash Kumar +1
UCF Institute of Artificial Intelligence, University of Central Florida (UCF)