Action representation plays a central role in discrete-token vision-language-action (VLA) learning but remains underexamined. Under conventional pose-increment representations, action tokens are sensitive to execution speed and dataset-specific normalization, potentially obscuring geometric structure shared across demonstrations and datasets. We introduce Direction-Scale Decomposition (DSD), an action representation that decomposes translation and rotation increments into direction and scale components before tokenization. DSD isolates motion direction while retaining magnitudes in separate scale channels. We evaluate DSD with uniform binning (BIN) and BEAST, a B-spline-based tokenizer, in simulation and real-world manipulation under both single-dataset and mixed-dataset training. On LIBERO, DSD improves average success rates with both tokenizers. On SimplerEnv, DSD-BIN outperforms BIN by 10.3 percentage points in overall success rate under mixed-dataset training. Real-robot experiments further show gains both with and without robotics pretraining. These results support DSD as an effective action representation for discrete-token VLA models and suggest its potential to mitigate performance degradation when training on large and diverse dataset mixtures. Our project page with additional resources is available at https://vla-dsd.github.io/
Figures & tables
Fig. 3: Light-tracking setup and normalized trajectory comparisons. From left to right: robot setup, demonstrations collected at five execution speeds, and 10 evaluation rollouts each for BIN, DSD-BIN, BEAST, and DSD-BEAST. Gold dashed curves denote the reference trajectory, and cross markers indicate safety-guard terminations. BEAST and DSD-BEAST each have two terminated rollouts.
Method
Average
Best
DTW distance ( ×10−2 ) ↓
BIN
1.997
1.414
DSD-BIN
1.588
0.781
BEAST
2.743
1.222
DSD-BEAST
2.465
0.947
Score with 0.05 cutoff ↑
TABLE I: Average and best tracking performance over 10 rollouts. Best denotes minimum DTW distance or maximum score.
Method
P
Spatial
Object
Goal
Long
Avg.
BIN
–
83.4
93.4
91.8
85.2
88.5
DSD–BIN
–
93.6
96.8
95.0
83.8
92.3
BEAST
–
88.4
96.4
88.6
78.4
88.0
DSD–BEAST
–
92.4
97.2
89.4
88.6
91.9
π0.5
yes
98.0
97.6
96.6
82.0
93.6
π -FAST
yes
96.4
98.0
87.8
60.0
85.6
TABLE II: LIBERO success rates (%). “P” indicates robotics pretraining. The highest success rate in each column is shown in bold.
Fig. 4: Real-robot manipulation results. Top: task setups for cube stacking, cloth folding, tissue sweeping, and spoon replacement. Bottom: success rates over 30 trials per method per task, with the four-task average shown on the right. Hatched bars indicate robotics pretraining, and green brackets show gains from adding DSD in percentage points.
Method
Carrot
Spoon
Stack
Eggplant
Overall
RT-1-X
10.7
4.0
0.0
0.0
3.7
Octo-base
6.7
8.0
0.0
41.3
14.0
Octo-small
5.3
34.7
2.7
53.3
24.0
OpenVLA
0.0
0.0
0.0
0.0
0.0
π -FAST-ft
0.0
0.0
0.0
9.3
2.3
π0 -ft
8.0
6.7
6.7
14.7
9.0
TABLE III: SimplerEnv success rate (%). “-bridge” = fine-tuned on bridgeV2 only; “-mix” = co-trained on the 5-dataset mixture.
The Hong Kong Polytechnic University, HongKong SAR, China · Shanghai AI Laboratory, Shanghai, China · The University of Hong Kong, HongKong SAR, China +2