ActionCodec: What Makes for Good Action Tokenizers
Authors: Zibin Dong, Yicheng Liu, Shiduo Zhang, Baijun Ye, Yifu Yuan, Fei Ni, Jingjing Gong, Xipeng Qiu, +3 more
Organizations: Knowin AI (Work done during an internship) · Tsinghua University · Tianjin University · Fudan University · Shanghai Innovation Institute
Vision-Language-Action (VLA) models leveraging the native autoregressive paradigm of Vision-Language Models (VLMs) have demonstrated superior instruction-following and training efficiency. Central to this paradigm is action tokenization, yet its design has primarily focused on reconstruction fidelity, failing to address its direct impact on VLA optimization. Consequently, the fundamental question of \textit{what makes for good action tokenizers} remains unanswered. In this paper, we bridge this gap by establishing design principles specifically from the perspective of VLA optimization. We identify a set of best practices based on information-theoretic insights, including maximized temporal token overlap, minimized vocabulary redundancy, enhanced multimodal mutual information, and token independence. Guided by these principles, we introduce \textbf{ActionCodec}, a high-performance action tokenizer that significantly enhances both training efficiency and VLA performance across diverse simulation and real-world benchmarks. Notably, on LIBERO, a SmolVLM2-2.2B fine-tuned with ActionCodec achieves a 95.5% success rate without any robotics pre-training. With advanced architectural enhancements, this reaches 97.4%, representing a new SOTA for VLA models without robotics pre-training. We believe our established design principles, alongside the released model, will provide a clear roadmap for the community to develop more effective action tokenizers.
Figures & tables
Figure 1: ActionCodec provides a comprehensive analysis of VQ action tokenizers that directly impact VLA training and summarizes the best practices. When utilized for the fine-tuning of SmolVLM2-2.2B without additional architectural designs, ActionCodec achieves performance on LIBERO benchmark that far exceeds other tokenizers, particularly in terms of training efficiency.
Figure 2: Neural Network Architecture of ActionCodec. We employ a Perceiver-like transformer architecture due to its inherent flexibility, which facilitates the modeling of diverse token relations and supports the encoding of variable-length action sequences.
Figure 3: LIBERO-Goal results for different design choices. All VLA models are based on SmolVLM2-256M, following vocabulary expansion and full-parameter fine-tuning without additional architectural modifications. The suffix notations are defined as follows: acc (L1 error), tok (token budget), cb (vocabulary size), OR (overlap rate), SA (self-attention), Causal (SA w/ causal mask), and SP (training w/ soft-prompt).
Figure 4
Figure 6: Error propagation under token perturbation. We measure the L1 reconstruction error when noise is injected at specific positions during generating. VQ-Independent maintains a stable, lower error compared to SA and Causal architectures. This supports reducing architectural token coupling for prefix-error robustness.
Tokenizer
Token Budget
OR
Horizon
Latency (s)
Throughput (action/s)
SR (%)
Binning [ 12 ]
140.0±0.0
39%
20
4.9
4.1
3.5
Binning [ 12 ]
56.0±0.0
38%
8
2.6
3.1
53.4
String [ 13 ]
779.2±17.3
28%
20
33.8
0.6
20.4
String [ 13 ]
311.8±7.8
34%
8
12.7
0.6
49.6
VQVLA’s [ 36 ]
4.0±0.0
45%
5
0.7
7.7
60.5
MiniVLA’s [ 2 ]
7.0±0.0
40%
8
0.7
11.1
82.6
Table 2: Action Tokenization Performance and Efficiency. Comparison of token budget, overlap rate (OR), and inference metrics. All models are evaluated on the LIBERO benchmark.
Figure 7: Success Rates on SO100-ShapeSorter. We report success counts for Pick (solid) and Place (hatched) across 10 tasks. While recovery actions are inherently present in our task demonstrations due to the difficulty of slot insertion, only ActionCodec (with CT) successfully learns to reproduce these corrective strategies. The performance gap is most prominent in the Place phase, where co-training facilitates the execution of complex recovery behaviors that task-only models fail to capture.
Table 8
Figure 8: Comparison of recovery behaviors on SO100. After an initial missed shot, ActionCodec with co-training successfully executes a corrective adjustment to complete the insertion, whereas the variant without co-training stalls after a failure.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Cross-embodiment action transfer. To evaluate the transferability of ActionCodec, a 1-second WidowX action sequence sampled from BridgeData is encoded into action tokens and subsequently decoded into 1-second trajectories for LIBERO-Franka, DROID-Franka, and xArm. After deploying these sequences twice (totaling 2 seconds) and recording the resulting motions, all three robotic platforms exhibit highly consistent action patterns. This demonstrates that ActionCodec effectively captures hardware-agnostic action semantics.
Figure 10: All evaluation benchmarks. We evaluate VLA models on 4 LIBERO task suites, Bridge-WidowX tasks on SimplerEnv, SO100-ShapeSorter tasks, and xArm tabletop manipulation tasks.
OR (%)
Codebook utilization (%)
PPL
L1 error
SR at 500 steps (%)
26
100
2016
0.020
8.2
40
100
1583
0.029
12.4
70
100
1290
0.028
33.4
Appendix
Table 5: Codebook usage in the controlled OR experiment. All variants use n=16 and S=2048 . Utilization is measured on the validation set; PPL is the marginal token-usage perplexity exp(H(c)) , with entropy in nats. Reconstruction errors are from Figure 3 .
Figure 11: Comparison of Perceiver architectures for Residual Grammar: (left) Independent (cross-attention only), (middle) SA (with self-attention), and (right) Causal (with causal self-attention)
Hyperparameters
Values
GPUs
8 × NVIDIA RTX 3090 (24 GB VRAM)
Learning Rate
2×10−4 peak (1k steps linear warmup, 30k steps cosine decay to 2×10−5 )
Global Batch Size
128 (16 per GPU)
Training Steps
30k (LIBERO-Goal only)
Input Modalities
1 × Third-person camera, 1 × Wrist camera (concatenated)
Input Image Size
224 × 448
Appendix
Table 6: VLA Training Hyperparameters. Detailed configurations for the validation experiments on LIBERO-Goal.
Hyperparameters
Values
GPUs
4 × NVIDIA A100 (80 GB VRAM)
Learning Rate
10−4 peak (1k steps linear warmup, 30k steps cosine decay to 10−5 )
Global Batch Size
128 (32 per GPU)
Training Steps
30k (applied to all four LIBERO task suites)
Input Modalities
1 × Third-person camera, 1 × Wrist camera (concatenated)
Input Image Size
224 × 448
Appendix
Table 7: VLA Training Hyperparameters. Detailed configurations for the tokenizer comparison experiments on LIBERO.
Hyperparameters
Values
GPUs
8 × NVIDIA H100 (94 GB VRAM)
Learning Rate
2×10−5
Global Batch Size
128 (16 per GPU)
Training Epochs
4
Input Modalities
1 × Third-person camera
Input Image Size
224 × 224
Appendix
Table 8: VLA Training Hyperparameters. Detailed configurations for the Simpler-WidowX experiments.
Hyperparameters
Values
GPUs
4 × NVIDIA A100 (80 GB VRAM)
Learning Rate
5×10−5 peak (1k steps linear warmup, 30k steps cosine decay to 5×10−6 )
Global Batch Size
64 (16 per GPU)
Training Steps
30k
Input Modalities
SO100: 1 × Third-person camera, 1 × Wrist camera (concatenated), xArm: 1 × Third-person camera
Input Image Size
224 × 448 for SO100, 224 × 224 for xArm
Appendix
Table 9: VLA Training Hyperparameters. Detailed configurations for the real-world experiments on SO100 and xArm.
Figure 12: VLA paradigm architectures.
Figure 13: Prompt template for VLM fine-tuning.
Figure 14: Learning curves of ActionCodec with and without pre-training. Compared to random initialization, initializing with pre-trained ActionCodec weights and fine-tuning on novel robotic data results in significantly lower L2 loss (higher reconstruction accuracy) and a higher Overlap Rate (OR). These results demonstrate that the pre-trained ActionCodec possesses a unified and semantically rich latent space, enabling rapid adaptation to new robotic action patterns.
Figure 15: Illustration of the Soft-Prompt integration within the Perceiver architecture. Learnable embeddings represent embodiment IDs, while fourier embeddings provide temporal grounding for varying control frequencies.
Figure 16: Procedure of RVQ post-training.
Method
Simpler-WidowX FT
LIBERO FT
DROID ZS
π0 -FAST
32.1
85.5
–
π0.5
57.1
96.8
57.5
MolmoAct2-DROID
–
–
52.0
ActionCodec (Qwen3.5-2B)
87.7
98.2
82.5
Appendix
Table 10: Large-scale VLA pre-training and transfer. Benchmark-average success rates (%). FT denotes downstream task fine-tuning; ZS denotes zero-shot evaluation. A dash denotes an unreported result.
Tokenizer
Encoding time per chunk (ms)
FAST
0.11
ActionCodec
0.09
Appendix
Table 11: Action encoding cost. Per-chunk encoding time at batch size 1024. These measurements concern tokenizer encoding, separately from policy generation latency in Table 2 .
The Hong Kong Polytechnic University, HongKong SAR, China · Shanghai AI Laboratory, Shanghai, China · The University of Hong Kong, HongKong SAR, China +2