Organizations: The Hong Kong University of Science and Technology, Hong Kong SAR, China. · Astribot, China. · The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, Guangdong, China.
Comparing current and past observations helps robots understand scene changes and select subsequent actions during manipulation. However, comparing visual features at the same image location can mix different scene content when objects or the camera move. We introduce CoRe-WAM, a world-action model that incorporates correspondence-aligned visual changes through a parameter-efficient temporal interface. Its TraceDelta module uses correspondences from a frozen tracking model to transport historical visual features to current locations before computing signed differences in a shared pretrained feature space. Correspondence thus determines which historical content is compared with the present, rather than entering the policy as a separate trajectory representation. A lightweight adapter converts these differences into validity-gated residuals that supplement current visual conditioning, allowing the policy to use recent changes alongside current-scene information. Built on Motus, CoRe-WAM keeps the pretrained backbone weights frozen and optimizes 1.59 million parameters. With a 5,000-update adaptation budget, CoRe-WAM achieves 92.22% clean success across 50 RoboTwin 2.0 tasks, 3.56 percentage points above Motus; on randomized evaluation, it achieves 89.60% success, a 2.58-point gain. Integrating TraceDelta into a StarVLA-based policy improves clean success from 58.10% to 67.62%, supporting transfer of the temporal interface beyond Motus.
Figures & tables
Fig. 1: Why align history before differencing? The current point p and historical point p′ share the same image coordinate but correspond to different scene content. Correspondence instead links p to its historical counterpart q , enabling feature differencing between matched regions. Images and correspondences are illustrative.
Fig. 2: CoRe-WAM retains current-observation conditioning while adding correspondence-aligned visual change. (a) A shared frozen vision-language model encodes current and historical observations. Current features pass through the frozen projector; TraceDelta uses both feature streams, and its residual is added to the projected current tokens. Instruction and state condition the Motus understanding, video, and action experts, whose joint attention generates an action chunk. (b) Frozen Trace Anything maps current point p to historical point q and transports its feature. TraceDelta subtracts this aligned feature from the current feature, adapts the signed difference, masks invalid entries to zero, and updates only head-camera tokens. Images and correspondences are schematic.
Method
Clean
Rand.
ΔC
ΔR
π0.5
82.74
76.76
—
—
X-VLA
72.88
72.84
—
—
Track4Action [ 33 ]
80.44
81.48
—
—
Fast-WAM
91.88
91.78
—
—
FastWAM-Joint
90.84
90.32
—
—
Motus
88.66
87.02
—
—
TABLE I: RoboTwin 2.0 success (%) and reported gains (pp). Bold/underline mark the highest/second-highest SR in each column. Shading identifies our method. ΔC / ΔR are relative to Motus, except for 4D-WAM, which uses FastWAM-Joint. Protocols follow Section IV-A .
Selected task
π0.5
X-VLA
FastWAM-Joint
Motus
CoRe-WAM
Δ
Move Can Pot
51
89
97
34
84
+50
Scan Object
72
14
92
67
86
+19
Handover Mic
98
0
100
78
94
+16
Put Bottles Dustbin
84
74
93
81
94
+13
Place A2B Left
87
48
96
88
98
+10
Place Bread Skillet
85
77
90
86
96
+10
TABLE II: Task-level clean success rates (%) on RoboTwin 2.0. The 18 tasks with gains of at least 5 pp over the base policy are ordered by Δ . Bold/underline mark the highest/second-highest displayed score, including ties. Shaded columns show CoRe-WAM and its gain over Motus. TraceDelta clean results are aggregated across two seeds following Section IV-A ; the last row equally weights all 50 tasks.
Fig. 3: Four RoboTwin examples show how task-relevant scene content evolves during manipulation. From top to bottom, the rows depict hammering, object handover, opening a laptop, and stacking three objects. Each row contains nine uniformly sampled head-camera frames; the numbers above them are original frame indices. The strips visualize the temporal evidence available to TraceDelta, which compares corresponding content within a short local history before conditioning the next action chunk.
Policy
Clean
Randomized
StarVLA
58.10
10.60
StarVLA + TraceDelta
67.62
10.81
Gain (pp)
+9.52
+0.21
TABLE III: Transferability to StarVLA under the same clean-data adaptation budget. SR is in %; gains are relative to StarVLA.
Adaptation components
Success rate (%)
Variant
LoRA
Adapter
History
Align
Params
Clean
Rand.
Avg.
Motus (published)
×
×
×
×
—
88.66
87.02
87.84 (-3.07)
LoRA-only
✓
×
×
×
0.934
90.90
87.30
89.10 (-1.81)
Current-only
✓
✓
×
×
1.594
91.34
87.82
89.58 (-1.33)
Raw Delta
✓
✓
✓
×
1.594
90.16
87.66
88.91 (-2.00)
TraceDelta (ours)
✓
✓
✓
✓
1.594
92.22
89.60
90.91
TABLE IV: Component analysis on RoboTwin 2.0. All adapted variants update the action decoder; LoRA denotes additional action-side low-rank updates. Params: optimized parameters (M). SR is in %; Avg. equally weights Clean and Randomized. Parentheses show the numerical gap to the complete model’s reported Avg. (pp).
Task
LoRA-only
Current-only
Raw Delta
Beat Block Hammer
82
85
88
Handover Block
75
80
82
Open Laptop
86
92
90
Stack Blocks Three
68
71
73
All tasks (50)
87.30
87.82
87.66
TABLE V: Task-level randomized controls (%). The four tasks correspond to the manipulation types in Fig. 3 ; the final row covers all 50 tasks. Bold marks the highest score in each row.
Component
Parameters
Temporal adapter and residual scale
660,225
Action-side LoRA
917,504
Existing action decoder
16,398
Newly introduced
1,577,729
Total optimized
1,594,127
Share of policy parameters
≈ 0.02%
TABLE VI: Optimized parameter budget for the Motus-based model. Newly introduced parameters exclude the existing action decoder. Only approximately 0.02% of the Motus-based policy parameters are optimized; the pretrained backbone and correspondence estimator remain frozen.
Fig. 4: Physical deployment on Astribot S1. Left: onboard cameras, dual arms and grippers, and TraceDelta’s head-view input. Right: plush-toy placement and an OOD sequence placing a banana on a blue plate, then a strawberry in a pink bowl. Bottom: the policy observes three RGB views and robot state, predicts a 48-step action chunk, executes a short segment, and reobserves.
Task
Split
Motus
CoRe-WAM
Plush-toy placement
ID
20/30 (66.7)
28/30 (93.3)
Fruit placement
OOD
18/30 (60.0)
25/30 (83.3)
TABLE VII: Real-world success on Astribot S1 over 30 trials per task. Entries show successes/trials (SR, %). ID denotes in-distribution.