We present an action sequence transfer system that adaptively transfers user action sequences across different target spaces. Given an input action sequence from a source space and scene graph representations of both the source and target environments, our system predicts a corresponding action sequence in the target space by adapting to the spatial and object constraints of the new environment. To achieve this, we leverage multi-level representations of user activity to generalize actions at varying levels of abstraction. To demonstrate our system, we collect a new scene graph-based dataset derived from the Ego4D GoalStep dataset for evaluation. Results indicate that our system can generate valid action sequences even between spaces with drastically different object configurations.
Figures & tables
Fig. 1 : Action sequence transfer problem between two different spaces: (a) Action sequence in source space is modified to match the environments of (b) target space, so that the overarching core activity-the goal-of the original action sequence is preserved.
Fig. 2 : Our action sequence transfer system predicts core activity, middle level core activity property, and the final action sequence for the target space using a sequence of LLM chains. Each LLM prompt in a chain is written to focus on generation of a single representation format. We visualize the resulting target action sequence after dividing each target action into executable action primitives for animation.
Fig. 3 : Our dataset consists of two components: (1) a reformatted Ego4D Goal-Step dataset, and (2) a dataset containing initial scene graphs along with corresponding scene graph changes aligned with its sub-steps as actions. We select 100 pairs of Goal-Step sequences and their corresponding initial spatial scene graphs as the evaluation dataset. The remaining data are stored in a vector store and used for prompt augmentation during the core activity prediction step.
scenario count for spatial dataset annotation
Scenario
Cooking
Baking
Laundry
Talking
Eating
Total Time
All Dataset
166
32
28
9
5
44.0 hours
Evaluation
83
6
9
8
2
22.6 hours
TABLE I : Scenario count (overlap is allowed) statistics for the annotated spatial dataset and the evaluation dataset splits.
Fig. 4 : A two-step target scene augmentation procedure is used to generate unique source–target space test pairs for system input.
Fig. 5 : Our framework enables robot in target space to actively use alternative objects to achieve activity goal, when corresponding objects are not available in the target environment.
Fig. 7 : For systems without fine-grained activity representations, success rates rapidly decrease for smaller scene inclusion ratios.
Success Percentage for Filtered Samples
Inclusion Ratio Range
0–0.33
0.33–0.67
0.67–1.0
Total (success / filtered)
Direct
81.2
90.0
94.8
90.4 (292 / 323)
Goal (no RAG)
82.1
96.7
98.1
94.3 (297 / 315)
Goal (RAG)
85.9
96.3
97.5
94.5 (328 / 347)
Property (no RAG)
94.0
93.7
97.3
95.5 (299 / 313)
Property (RAG, Ours )
96.0
95.9
98.1
97.0 (318 / 328)
TABLE II : Percentage of successful transfer across inclusion ratio ranges
Fig. 8 : In visualization, target action sequence is converted to action primitives for animating robot agent in target space.
Action Count for a Successful Pair of Sequences
Direct
Goal
Goal-RAG
Prop.
Prop.-RAG
Source
10.7 (7.6)
10.4 (7.4)
10.4 (7.1)
10.3 (7.2)
10.3 (7.2)
Target
8.0 (6.2)
5.6 (5.0)
5.7 (4.8)
4.0 (3.2)
4.1 (2.8)
TABLE III : Mean and standard deviation of action counts in source action sequence and target action sequence for samples with successful action sequence transfer.
Despite the central role of action in embodied intelligence, learning transferable action representations from visual transitions remains a fundamental challenge, particularly when world models must generalize across embodiments under limited data. We argue that action is not merely an auxiliary conditioning signal, but a distinct representational factor that decouples the controllable change from embodiment-specific actuation. In this work, we propose SCAR, a joint inverse-forward dynamics framework for learning unified action representations across embodiments from visual transitions. Built on a pretrained generative backbone, SCAR uses an inverse dynamics model (IDM) to infer latent actions from latent observation pairs and a forward dynamics model (FDM) to predict future dynamics conditioned on them. To make the latent space transferable rather than a generic visual bottleneck, we regularize the latent action posterior toward a standard Gaussian prior to limit arbitrary visual encoding, and introduce adversarial invariance to suppress embodiment- and environment-specific nuisance factors. Experiments on the Procgen and Robotwin dataset show that the learned unified latent action representation serves as a stronger conditioning interface for world modeling than embodiment-specific raw actions, yielding improved cross-embodiment low-data adaptation and cross-task transfer. Taken together, these results suggest that action can be learned as a shared representation of controllable change across embodiments, providing an interface for more transferable and generalizable world models.
Hongjia Liu, Fan Feng, Minghao Fu +3
University of California, San Diego · KTH Royal Institute of Technology
General-purpose embodied manipulation hinges on a unified action representation that generalizes across embodiments and scales readily. Yet existing policies rely on embodiment-specific action spaces, making cross-embodiment demonstrations difficult to leverage at scale and limiting transfer to new embodiments and spatial variations. To this end, we introduce Universal Manipulation Representation (UMR), a unified action representation that enables zero-shot skill transfer from human demonstrations to heterogeneous robots. UMR decomposes manipulation into two functionally distinct yet geometrically linked components: embodiment-agnostic World Flow, which describes task-relevant object motion in the world frame, and Ego Trajectory, which represents end-effector motion relative to the current pose. We instantiate UMR as World--Ego Point VLA (WEPVLA), a compact 0.5B-parameter policy that learns in the unified geometric action space through a dual-stream Point Action Adapter and a unified Point Action Expert, with an SE(3) conjugation coupling the two components. To improve data efficiency, we complement UMR with a Data-Efficient Strategy (DES) that diversifies object configurations through stage-aware point-cloud editing while preserving demonstrated contact geometry. In simulation, WEPVLA achieves average success rates of 97.5% on LIBERO and 85.7% on the 10-task RLBench benchmark. In real-world experiments, a single policy trained on human demonstrations augmented by DES transfers zero-shot to diverse deployment conditions. With about 10 minutes of collected human demonstrations per task and no robot demonstrations, it achieves 91.7% average success across six evaluation settings, compared with 60.8% for HumanEgo. Code and additional materials are available at https://umr-wepvla.github.io/.
Song Liu, Linyi Li, Yanshun Zhao +11
University of Science and Technology of China, Hefei, China. · Suzhou Artificial Intelligence Laboratory, Suzhou, China.
Humans can effortlessly perceive spatial layouts, form cognitive representations, reason about spatial relations, and translate such reasoning into actions in everyday 3D environments. Although recent vision-language models (VLMs) have shown promising performance on observation-conditioned spatial perception and reasoning tasks, it remains unclear whether they can build coherent spatial understanding, act upon it, and refine their actions through multi-turn feedback. To study this problem, we introduce \textbf{SpatialAct}, a simulator-grounded benchmark for probing \textit{action-conditioned spatial reasoning} in 3D scenes. Starting from the most challenging setting, Multi-turn Interactive Refinement, we further design its decomposed counterpart, Single-step Error Detection and Fix, together with five fundamental spatial ability tasks to diagnose the underlying causes of model failures. Experiments reveal a clear reasoning-to-action gap: current VLMs can perform well on isolated spatial reasoning tasks, but struggle to maintain coherent spatial beliefs and produce reliable actions during multi-turn feedback, substantially underperforming humans. These results suggest that current VLM agents still lack robust spatial state tracking under action-induced environment changes, even when low-level control is abstracted away.
Tianhui Liu, Jie Feng, Zhiheng Zheng +6
1The Hong Kong University of Science and Technology (Guangzhou) · 2Zhongguancun Academy · 3Tsinghua University +1