Action representation plays a central role in discrete-token vision-language-action (VLA) learning but remains underexamined. Under conventional pose-increment representations, action tokens are sensitive to execution speed and dataset-specific normalization, potentially obscuring geometric structure shared across demonstrations and datasets. We introduce Direction-Scale Decomposition (DSD), an action representation that decomposes translation and rotation increments into direction and scale components before tokenization. DSD isolates motion direction while retaining magnitudes in separate scale channels. We evaluate DSD with uniform binning (BIN) and BEAST, a B-spline-based tokenizer, in simulation and real-world manipulation under both single-dataset and mixed-dataset training. On LIBERO, DSD improves average success rates with both tokenizers. On SimplerEnv, DSD-BIN outperforms BIN by 10.3 percentage points in overall success rate under mixed-dataset training. Real-robot experiments further show gains both with and without robotics pretraining. These results support DSD as an effective action representation for discrete-token VLA models and suggest its potential to mitigate performance degradation when training on large and diverse dataset mixtures. Our project page with additional resources is available at https://vla-dsd.github.io/
Figures & tables
Fig. 3: Light-tracking setup and normalized trajectory comparisons. From left to right: robot setup, demonstrations collected at five execution speeds, and 10 evaluation rollouts each for BIN, DSD-BIN, BEAST, and DSD-BEAST. Gold dashed curves denote the reference trajectory, and cross markers indicate safety-guard terminations. BEAST and DSD-BEAST each have two terminated rollouts.
Method
Average
Best
DTW distance ( ×10−2 ) ↓
BIN
1.997
1.414
DSD-BIN
1.588
0.781
BEAST
2.743
1.222
DSD-BEAST
2.465
0.947
Score with 0.05 cutoff ↑
TABLE I: Average and best tracking performance over 10 rollouts. Best denotes minimum DTW distance or maximum score.
Method
P
Spatial
Object
Goal
Long
Avg.
BIN
–
83.4
93.4
91.8
85.2
88.5
DSD–BIN
–
93.6
96.8
95.0
83.8
92.3
BEAST
–
88.4
96.4
88.6
78.4
88.0
DSD–BEAST
–
92.4
97.2
89.4
88.6
91.9
π0.5
yes
98.0
97.6
96.6
82.0
93.6
π -FAST
yes
96.4
98.0
87.8
60.0
85.6
TABLE II: LIBERO success rates (%). “P” indicates robotics pretraining. The highest success rate in each column is shown in bold.
Fig. 4: Real-robot manipulation results. Top: task setups for cube stacking, cloth folding, tissue sweeping, and spoon replacement. Bottom: success rates over 30 trials per method per task, with the four-task average shown on the right. Hatched bars indicate robotics pretraining, and green brackets show gains from adding DSD in percentage points.
Method
Carrot
Spoon
Stack
Eggplant
Overall
RT-1-X
10.7
4.0
0.0
0.0
3.7
Octo-base
6.7
8.0
0.0
41.3
14.0
Octo-small
5.3
34.7
2.7
53.3
24.0
OpenVLA
0.0
0.0
0.0
0.0
0.0
π -FAST-ft
0.0
0.0
0.0
9.3
2.3
π0 -ft
8.0
6.7
6.7
14.7
9.0
TABLE III: SimplerEnv success rate (%). “-bridge” = fine-tuned on bridgeV2 only; “-mix” = co-trained on the 5-dataset mixture.
Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.
Shijie Lian, Bin Yu, Zhaolong Shen +5
Huazhong University of Science and Technology · Zhongguancun Academy · DeepCybo +5
Discrete action tokenization provides a compact interface for autoregressive VLA policies, but accurately recovering continuous robot actions from discrete codes remains challenging. Existing tokenizers typically map each discrete code to a fixed continuous action prototype, ignoring the robot's current proprioceptive state. This limitation is particularly pronounced in manipulation, where the same action token may require different continuous controls under different joint configurations, object poses, and contact conditions. We therefore propose SA-VLA, a state-aware action tokenizer that conditions action decoding on robot state. We study two state-injection mechanisms for VQ-based action tokenization: cross-attention between state and action features, and a lightweight state adapter that predicts action-wise modulation factors for state-conditioned action modulation and reconstruction. The adapter formulation expands the effective support of a finite codebook by allowing each discrete token to represent a family of state-dependent continuous actions, while preserving the efficiency and compatibility of discrete action modeling. Integrated into an LLM-based VLA policy, SA-VLA supports both autoregressive and parallel action-token decoding with minimal changes to the model interface. On 12 RoboTwin manipulation tasks, SA-VLA improves the average success rate from 0.29 to 0.56 over the strongest tokenizer baseline. In zero-shot sim-to-real experiments on three real-world tasks, it further improves average success from 0.15 to 0.33 over the strongest tokenizer baseline. These results demonstrate that state-conditioned action decoding is a simple and effective mechanism for reducing the compression gap in discrete VLA policies.
Tengyue Jiang, Chunpu Xu, Jiayue Kang +1
2East China University of Science and Technology · 3Hong Kong Polytechnic University · 4Xi’an University of Electronic Science and Technology +1
Recent advancements have successfully adapted autoregressive language models to process multimodal signals, such as images and actions. Since raw action signals are continuous, effective tokenization is essential to map high-dimensional inputs into compact discrete tokens for autoregressive processing. However, existing discrete action tokenizers often suffer from high reconstruction loss, failing to preserve the fine-grained dynamics required for precise control. This "discretization bottleneck" significantly limits the performance ceiling of downstream Vision-Language-Action (VLA) models. To address this, we propose M2Tok, a Multi-head Multi-codebook Action Tokenizer designed to minimize reconstruction error and enhance policy performance. Our approach introduces two key structural innovations: (1) we decompose the latent action features into multiple heads, enabling the model to implicitly align specific heads with distinct action dimensions; (2) we assign independent codebooks to each head for quantization. By leveraging the combinatorial nature of multiple codebooks, we significantly expand the representational expressivity of the tokenizer, leading to substantially lower reconstruction loss compared to previous methods. We evaluate the M2Tok-based VLA on the RoboTwin, Simpler-Env, and 3 zero-shot real-world tasks. Experimental results demonstrate our method not only achieves superior reconstruction fidelity but also significantly boosts the success rate of VLA models. Comprehensive ablation studies further confirm the effectiveness of the multi-head and multi-codebook mechanisms. Code is available at https://github.com/cpaaax/M2Tok.
Chunpu Xu, Zhixuan Liang, Yuhao Zhang +6
The Hong Kong Polytechnic University, HongKong SAR, China · Shanghai AI Laboratory, Shanghai, China · The University of Hong Kong, HongKong SAR, China +2