Learning how actions change scene geometry can provide complementary supervision for goal-directed manipulation. We introduce Spatial Language Modeling, which represents scene contours, goals, action targets, and future states with a shared vocabulary of discrete coordinates and semantic tokens. A task-specific grammar organizes these elements into spatial sequences, allowing one autoregressive Transformer to learn action generation and action-conditioned state prediction through a common next-token objective. We train the model from scratch using random-play transition pretraining followed by joint action and state training on expert demonstrations. During pretraining, recorded action coordinates condition subsequent state predictions and are excluded from the prediction loss. During control, the model decodes only executable action targets and updates its history with newly observed states. We evaluate the approach on Push-T in simulation and on a real robot. The model achieves competitive simulation performance and higher task success and target coverage than the evaluated real-robot policy baselines. Training ablations show improved control with joint action and state sequences, with further gains from random-play pretraining. Given supplied action trajectories, the same model also predicts successive scene states, capturing the geometric effects of pushing.
Figures & tables
Figure 1: Spatial Language Modeling for Push-T. Scene geometry, goals, pusher actions, and optional future states are represented as structured coordinate-token sequences. The model generates actions for closed-loop execution and learns state transitions through future-state supervision.
Figure 2
Figure 3: SLS completion for action-conditioned state prediction (left) and action-only control decoding (right). Generated action blocks are parsed into pusher target coordinates.
Figure 4
Figure 5: Future-state prediction on Push-T. Given a manually annotated action trajectory, the model predicts successive states using its own previous predictions as context. Reference states are obtained by executing the same actions in the simulator. Subfigures (a) and (b) show two example rollouts, where red pixels indicate the prediction-error regions of the T-block on the rasterized workspace. Subfigure (c) reports the average state prediction accuracy at each successive step (frame) over the test rollouts, where dashed lines denote coverage and solid lines denote IoU.
Figure 6: For each method, images are averaged over the test episodes at their respective maximum-coverage frames. This visualization summarizes alignment across episodes.