ActionCodec: What Makes for Good Action Tokenizers
Authors: Zibin Dong, Yicheng Liu, Shiduo Zhang, Baijun Ye, Yifu Yuan, Fei Ni, Jingjing Gong, Xipeng Qiu, +3 more
Organizations: Knowin AI (Work done during an internship) · Tsinghua University · Tianjin University · Fudan University · Shanghai Innovation Institute
Vision-Language-Action (VLA) models leveraging the native autoregressive paradigm of Vision-Language Models (VLMs) have demonstrated superior instruction-following and training efficiency. Central to this paradigm is action tokenization, yet its design has primarily focused on reconstruction fidelity, failing to address its direct impact on VLA optimization. Consequently, the fundamental question of \textit{what makes for good action tokenizers} remains unanswered. In this paper, we bridge this gap by establishing design principles specifically from the perspective of VLA optimization. We identify a set of best practices based on information-theoretic insights, including maximized temporal token overlap, minimized vocabulary redundancy, enhanced multimodal mutual information, and token independence. Guided by these principles, we introduce \textbf{ActionCodec}, a high-performance action tokenizer that significantly enhances both training efficiency and VLA performance across diverse simulation and real-world benchmarks. Notably, on LIBERO, a SmolVLM2-2.2B fine-tuned with ActionCodec achieves a 95.5% success rate without any robotics pre-training. With advanced architectural enhancements, this reaches 97.4%, representing a new SOTA for VLA models without robotics pre-training. We believe our established design principles, alongside the released model, will provide a clear roadmap for the community to develop more effective action tokenizers.
Figures & tables
Figure 1: ActionCodec provides a comprehensive analysis of VQ action tokenizers that directly impact VLA training and summarizes the best practices. When utilized for the fine-tuning of SmolVLM2-2.2B without additional architectural designs, ActionCodec achieves performance on LIBERO benchmark that far exceeds other tokenizers, particularly in terms of training efficiency.
Figure 2: Neural Network Architecture of ActionCodec. We employ a Perceiver-like transformer architecture due to its inherent flexibility, which facilitates the modeling of diverse token relations and supports the encoding of variable-length action sequences.
Figure 3: LIBERO-Goal results for different design choices. All VLA models are based on SmolVLM2-256M, following vocabulary expansion and full-parameter fine-tuning without additional architectural modifications. The suffix notations are defined as follows: acc (L1 error), tok (token budget), cb (vocabulary size), OR (overlap rate), SA (self-attention), Causal (SA w/ causal mask), and SP (training w/ soft-prompt).
Figure 4
Figure 6: Error propagation under token perturbation. We measure the L1 reconstruction error when noise is injected at specific positions during generating. VQ-Independent maintains a stable, lower error compared to SA and Causal architectures. This supports reducing architectural token coupling for prefix-error robustness.
Tokenizer
Token Budget
OR
Horizon
Latency (s)
Throughput (action/s)
SR (%)
Binning [ 12 ]
140.0±0.0
39%
20
4.9
4.1
3.5
Binning [ 12 ]
56.0±0.0
38%
8
2.6
3.1
53.4
String [ 13 ]
779.2±17.3
28%
20
33.8
0.6
20.4
String [ 13 ]
311.8±7.8
34%
8
12.7
0.6
49.6
VQVLA’s [ 36 ]
4.0±0.0
45%
5
0.7
7.7
60.5
MiniVLA’s [ 2 ]
7.0±0.0
40%
8
0.7
11.1
82.6
Table 2: Action Tokenization Performance and Efficiency. Comparison of token budget, overlap rate (OR), and inference metrics. All models are evaluated on the LIBERO benchmark.
Figure 7: Success Rates on SO100-ShapeSorter. We report success counts for Pick (solid) and Place (hatched) across 10 tasks. While recovery actions are inherently present in our task demonstrations due to the difficulty of slot insertion, only ActionCodec (with CT) successfully learns to reproduce these corrective strategies. The performance gap is most prominent in the Place phase, where co-training facilitates the execution of complex recovery behaviors that task-only models fail to capture.
Table 8
Figure 8: Comparison of recovery behaviors on SO100. After an initial missed shot, ActionCodec with co-training successfully executes a corrective adjustment to complete the insertion, whereas the variant without co-training stalls after a failure.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Cross-embodiment action transfer. To evaluate the transferability of ActionCodec, a 1-second WidowX action sequence sampled from BridgeData is encoded into action tokens and subsequently decoded into 1-second trajectories for LIBERO-Franka, DROID-Franka, and xArm. After deploying these sequences twice (totaling 2 seconds) and recording the resulting motions, all three robotic platforms exhibit highly consistent action patterns. This demonstrates that ActionCodec effectively captures hardware-agnostic action semantics.
Figure 10: All evaluation benchmarks. We evaluate VLA models on 4 LIBERO task suites, Bridge-WidowX tasks on SimplerEnv, SO100-ShapeSorter tasks, and xArm tabletop manipulation tasks.
OR (%)
Codebook utilization (%)
PPL
L1 error
SR at 500 steps (%)
26
100
2016
0.020
8.2
40
100
1583
0.029
12.4
70
100
1290
0.028
33.4
Appendix
Table 5: Codebook usage in the controlled OR experiment. All variants use n=16 and S=2048 . Utilization is measured on the validation set; PPL is the marginal token-usage perplexity exp(H(c)) , with entropy in nats. Reconstruction errors are from Figure 3 .
Figure 11: Comparison of Perceiver architectures for Residual Grammar: (left) Independent (cross-attention only), (middle) SA (with self-attention), and (right) Causal (with causal self-attention)
Hyperparameters
Values
GPUs
8 × NVIDIA RTX 3090 (24 GB VRAM)
Learning Rate
2×10−4 peak (1k steps linear warmup, 30k steps cosine decay to 2×10−5 )
Global Batch Size
128 (16 per GPU)
Training Steps
30k (LIBERO-Goal only)
Input Modalities
1 × Third-person camera, 1 × Wrist camera (concatenated)
Input Image Size
224 × 448
Appendix
Table 6: VLA Training Hyperparameters. Detailed configurations for the validation experiments on LIBERO-Goal.
Hyperparameters
Values
GPUs
4 × NVIDIA A100 (80 GB VRAM)
Learning Rate
10−4 peak (1k steps linear warmup, 30k steps cosine decay to 10−5 )
Global Batch Size
128 (32 per GPU)
Training Steps
30k (applied to all four LIBERO task suites)
Input Modalities
1 × Third-person camera, 1 × Wrist camera (concatenated)
Input Image Size
224 × 448
Appendix
Table 7: VLA Training Hyperparameters. Detailed configurations for the tokenizer comparison experiments on LIBERO.
Hyperparameters
Values
GPUs
8 × NVIDIA H100 (94 GB VRAM)
Learning Rate
2×10−5
Global Batch Size
128 (16 per GPU)
Training Epochs
4
Input Modalities
1 × Third-person camera
Input Image Size
224 × 224
Appendix
Table 8: VLA Training Hyperparameters. Detailed configurations for the Simpler-WidowX experiments.
Hyperparameters
Values
GPUs
4 × NVIDIA A100 (80 GB VRAM)
Learning Rate
5×10−5 peak (1k steps linear warmup, 30k steps cosine decay to 5×10−6 )
Global Batch Size
64 (16 per GPU)
Training Steps
30k
Input Modalities
SO100: 1 × Third-person camera, 1 × Wrist camera (concatenated), xArm: 1 × Third-person camera
Input Image Size
224 × 448 for SO100, 224 × 224 for xArm
Appendix
Table 9: VLA Training Hyperparameters. Detailed configurations for the real-world experiments on SO100 and xArm.
Figure 12: VLA paradigm architectures.
Figure 13: Prompt template for VLM fine-tuning.
Figure 14: Learning curves of ActionCodec with and without pre-training. Compared to random initialization, initializing with pre-trained ActionCodec weights and fine-tuning on novel robotic data results in significantly lower L2 loss (higher reconstruction accuracy) and a higher Overlap Rate (OR). These results demonstrate that the pre-trained ActionCodec possesses a unified and semantically rich latent space, enabling rapid adaptation to new robotic action patterns.
Figure 15: Illustration of the Soft-Prompt integration within the Perceiver architecture. Learnable embeddings represent embodiment IDs, while fourier embeddings provide temporal grounding for varying control frequencies.
Figure 16: Procedure of RVQ post-training.
Method
Simpler-WidowX FT
LIBERO FT
DROID ZS
π0 -FAST
32.1
85.5
–
π0.5
57.1
96.8
57.5
MolmoAct2-DROID
–
–
52.0
ActionCodec (Qwen3.5-2B)
87.7
98.2
82.5
Appendix
Table 10: Large-scale VLA pre-training and transfer. Benchmark-average success rates (%). FT denotes downstream task fine-tuning; ZS denotes zero-shot evaluation. A dash denotes an unreported result.
Tokenizer
Encoding time per chunk (ms)
FAST
0.11
ActionCodec
0.09
Appendix
Table 11: Action encoding cost. Per-chunk encoding time at batch size 1024. These measurements concern tokenizer encoding, separately from policy generation latency in Table 2 .
Recent advancements have successfully adapted autoregressive language models to process multimodal signals, such as images and actions. Since raw action signals are continuous, effective tokenization is essential to map high-dimensional inputs into compact discrete tokens for autoregressive processing. However, existing discrete action tokenizers often suffer from high reconstruction loss, failing to preserve the fine-grained dynamics required for precise control. This "discretization bottleneck" significantly limits the performance ceiling of downstream Vision-Language-Action (VLA) models. To address this, we propose M2Tok, a Multi-head Multi-codebook Action Tokenizer designed to minimize reconstruction error and enhance policy performance. Our approach introduces two key structural innovations: (1) we decompose the latent action features into multiple heads, enabling the model to implicitly align specific heads with distinct action dimensions; (2) we assign independent codebooks to each head for quantization. By leveraging the combinatorial nature of multiple codebooks, we significantly expand the representational expressivity of the tokenizer, leading to substantially lower reconstruction loss compared to previous methods. We evaluate the M2Tok-based VLA on the RoboTwin, Simpler-Env, and 3 zero-shot real-world tasks. Experimental results demonstrate our method not only achieves superior reconstruction fidelity but also significantly boosts the success rate of VLA models. Comprehensive ablation studies further confirm the effectiveness of the multi-head and multi-codebook mechanisms. Code is available at https://github.com/cpaaax/M2Tok.
Chunpu Xu, Zhixuan Liang, Yuhao Zhang +6
The Hong Kong Polytechnic University, HongKong SAR, China · Shanghai AI Laboratory, Shanghai, China · The University of Hong Kong, HongKong SAR, China +2
Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.
Shijie Lian, Bin Yu, Zhaolong Shen +5
Huazhong University of Science and Technology · Zhongguancun Academy · DeepCybo +5
Discrete action tokenization provides a compact interface for autoregressive VLA policies, but accurately recovering continuous robot actions from discrete codes remains challenging. Existing tokenizers typically map each discrete code to a fixed continuous action prototype, ignoring the robot's current proprioceptive state. This limitation is particularly pronounced in manipulation, where the same action token may require different continuous controls under different joint configurations, object poses, and contact conditions. We therefore propose SA-VLA, a state-aware action tokenizer that conditions action decoding on robot state. We study two state-injection mechanisms for VQ-based action tokenization: cross-attention between state and action features, and a lightweight state adapter that predicts action-wise modulation factors for state-conditioned action modulation and reconstruction. The adapter formulation expands the effective support of a finite codebook by allowing each discrete token to represent a family of state-dependent continuous actions, while preserving the efficiency and compatibility of discrete action modeling. Integrated into an LLM-based VLA policy, SA-VLA supports both autoregressive and parallel action-token decoding with minimal changes to the model interface. On 12 RoboTwin manipulation tasks, SA-VLA improves the average success rate from 0.29 to 0.56 over the strongest tokenizer baseline. In zero-shot sim-to-real experiments on three real-world tasks, it further improves average success from 0.15 to 0.33 over the strongest tokenizer baseline. These results demonstrate that state-conditioned action decoding is a simple and effective mechanism for reducing the compression gap in discrete VLA policies.
Tengyue Jiang, Chunpu Xu, Jiayue Kang +1
2East China University of Science and Technology · 3Hong Kong Polytechnic University · 4Xi’an University of Electronic Science and Technology +1