Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action tokenizer that reframes tokenization as a causally structured generative process. CATok introduces a conditional annealing mechanism that extracts action tokens by progressively annealing a flow-matching process: each token is conditioned on all preceding tokens and encodes the residual reconstruction signal at a specific noise level, establishing a coarse-to-fine causal token space whose generative semantics are structurally aligned with autoregressive modeling. A token-conditioned flow-matching decoder built on Multimodal Diffusion Transformer (MMDiT) reconstructs continuous action chunks from these discrete tokens with the precision of hybrid diffusion-head architectures. This discrete bottleneck enforces knowledge insulation by design, cleanly separating high-level semantic reasoning from low-level motor execution without requiring explicit attention masking. Extensive evaluations across three simulation benchmarks and real-world robotic manipulation tasks demonstrate that CATok consistently surpasses existing tokenization methods in both reconstruction fidelity-compression tradeoff and inference efficiency, while improving VLA task success rate and training efficiency, establishing a high-performance, scalable foundation for purely autoregressive VLA systems.
Figures & tables
Figure 1 : Flow-Matching to Action Tokens. CATok grounds discrete action tokens in the continuous flow matching trajectory. Each token is learned as a stage-wise information increment, imbuing the sequence with a natural coarse-to-fine hierarchy. By construction, these increments follow the temporal evolution of the flow, ensuring a causally-ordered structure that seamlessly aligns with autoregressive generation.
Figure 2 : An Overview of CATok . (a) CATok Pipeline. CATok encodes an action chunk with a Dual-Stream Encoder into latent queries, discretizes them via a Bottleneck VQ , and reconstructs actions using a Flow-Matching Decoder with Conditional Annealing . (b) Conditional Annealing Mechanism progressively masks the first κ(t) token embeddings according to flow matching timestep t: darker colors indicate tokens earlier in the sequence, and lighter colors indicate later tokens, assigning each token to a different stage of the flow trajectory.
Figure 3 : An Overview of VLA- CATok . VLA- CATok feeds visual, language, and proprioceptive tokens into a VLM backbone to autoregressively predict causal action tokens. These tokens are mapped via the learned codebook and processed by a frozen MMDiT flow-matching decoder to produce the final action in a single denoising step. Discrete tokens serve as the sole interface, naturally insulating the VLM’s pretrained knowledge from action-specific gradients.
Figure 4 : Real-world scene layouts for the three task suites. From left to right: Pick-Spatial , testing spatial generalization by varying cup and plate positions; Pick-Color , testing color generalization with different unseen cup colors; and Stack-Long , testing long-horizon manipulation under different spatial layouts.
LIBERO
SimplerEnv
RoboTwin 2.0
Method
Spatial
Object
Goal
Long
Avg.
Spoon
Carrot
Stack
Eggplant
Avg.
Clean
Randomized
BIN
0.586
0.878
0.680
0.604
0.687
0.542
0.333
0.208
0.708
0.448
0.213
0.221
FAST
0.960
0.998
0.962
0.901
0.955
0.500
0.375
0.375
0.625
0.469
0.478
0.478
OAT
0.428
0.876
0.704
0.276
0.571
0.417
0.208
0.167
0.667
0.365
0.229
0.233
CATok
0.978
0.994
0.954
0.910
0.959
0.458
0.417
0.542
0.542
0.490
0.489
0.531
Table 1 : Comparison on robotic manipulation simulation benchmarks. CATok achieves the best overall performance across all benchmarks, with particularly significant improvements on long-horizon tasks such as Long in LIBERO and StackGreenCubeOnYellowCube in SimplerEnv, demonstrating the effectiveness of causal action tokenization for modeling long-range temporal dependencies.
Figure 5 : (a) Training efficiency on the SimplerEnv benchmark. Curves are Gaussian-smoothed to reduce evaluation stochasticity without changing the relative ranking of different methods. Notably, CATok reaches the best baseline performance using only around 50% of the training steps. (b) VLA inference latency. By jointly considering autoregressive generation and decoding overhead, CATok achieves near-minimal end-to-end inference latency while attaining the best success rate among all methods.
Method
Pick-Spatial
Pick-Color
Stack-Long
FAST
0.450
0.300
0.275
CATok
0.650
0.500
0.475
Table 2 : Real-world manipulation results. CATok consistently outperforms FAST across spatial, color, and long-horizon generalization tasks.
Figure 6 : CATok Produces Causal Tokens with Generative Semantics . (a) Causal Ordering CATok ’s entropy decreases in the forward order ( CATOK fwd ) and increases in the reverse order ( CATOK rev ), while FAST and Bin show weaker positional structure. (b) Token embedding geometry. T-SNE map shows slot-dependent organization, progressing from smooth early-token regions to compact late-token clusters. (c) Prefix reconstruction. Increasing token prefixes produce structured trajectory updates, with dashed curves showing incremental contributions.
Figure 7 : Intrinsic properties of action tokenizers. (a–c) Fidelity vs. compression efficiency. VRR measures reconstruction quality; CR measures compression ratio; VRR × CR jointly captures the fidelity–compression tradeoff. CATok 8 achieves the highest VRR × CR ( 16.85 ), outperforming all baselines at the same compression level. (d–e) Inference efficiency. Encode and decode latencies are shown on a log scale. The subscript k in OAT k and CATok k denotes the number of tokens.
Table 4: Controlled ablation of conditional annealing, random masking, and decoder architecture.
Figure 8 : Effect of the number of action tokens on downstream VLA performance on the LIBERO benchmark. Performance improves from 8 to 16 tokens but degrades with larger token budgets, with 16 tokens achieving the best overall performance.
Figure 9 : Causality validation via native-decoder interventions. Both donor-swap (a) and single-token removal (b) exhibit a clear position-dependent structure: early tokens have stronger influence on coarse action components and earlier generation stages, whereas later tokens increasingly affect fine-grained action details, confirming that the learned token ordering corresponds to meaningful causal roles rather than an arbitrary positional ordering.
Discrete action tokenization provides a compact interface for autoregressive VLA policies, but accurately recovering continuous robot actions from discrete codes remains challenging. Existing tokenizers typically map each discrete code to a fixed continuous action prototype, ignoring the robot's current proprioceptive state. This limitation is particularly pronounced in manipulation, where the same action token may require different continuous controls under different joint configurations, object poses, and contact conditions. We therefore propose SA-VLA, a state-aware action tokenizer that conditions action decoding on robot state. We study two state-injection mechanisms for VQ-based action tokenization: cross-attention between state and action features, and a lightweight state adapter that predicts action-wise modulation factors for state-conditioned action modulation and reconstruction. The adapter formulation expands the effective support of a finite codebook by allowing each discrete token to represent a family of state-dependent continuous actions, while preserving the efficiency and compatibility of discrete action modeling. Integrated into an LLM-based VLA policy, SA-VLA supports both autoregressive and parallel action-token decoding with minimal changes to the model interface. On 12 RoboTwin manipulation tasks, SA-VLA improves the average success rate from 0.29 to 0.56 over the strongest tokenizer baseline. In zero-shot sim-to-real experiments on three real-world tasks, it further improves average success from 0.15 to 0.33 over the strongest tokenizer baseline. These results demonstrate that state-conditioned action decoding is a simple and effective mechanism for reducing the compression gap in discrete VLA policies.
Tengyue Jiang, Chunpu Xu, Jiayue Kang +1
2East China University of Science and Technology · 3Hong Kong Polytechnic University · 4Xi’an University of Electronic Science and Technology +1
Recent advancements have successfully adapted autoregressive language models to process multimodal signals, such as images and actions. Since raw action signals are continuous, effective tokenization is essential to map high-dimensional inputs into compact discrete tokens for autoregressive processing. However, existing discrete action tokenizers often suffer from high reconstruction loss, failing to preserve the fine-grained dynamics required for precise control. This "discretization bottleneck" significantly limits the performance ceiling of downstream Vision-Language-Action (VLA) models. To address this, we propose M2Tok, a Multi-head Multi-codebook Action Tokenizer designed to minimize reconstruction error and enhance policy performance. Our approach introduces two key structural innovations: (1) we decompose the latent action features into multiple heads, enabling the model to implicitly align specific heads with distinct action dimensions; (2) we assign independent codebooks to each head for quantization. By leveraging the combinatorial nature of multiple codebooks, we significantly expand the representational expressivity of the tokenizer, leading to substantially lower reconstruction loss compared to previous methods. We evaluate the M2Tok-based VLA on the RoboTwin, Simpler-Env, and 3 zero-shot real-world tasks. Experimental results demonstrate our method not only achieves superior reconstruction fidelity but also significantly boosts the success rate of VLA models. Comprehensive ablation studies further confirm the effectiveness of the multi-head and multi-codebook mechanisms. Code is available at https://github.com/cpaaax/M2Tok.
Chunpu Xu, Zhixuan Liang, Yuhao Zhang +6
The Hong Kong Polytechnic University, HongKong SAR, China · Shanghai AI Laboratory, Shanghai, China · The University of Hong Kong, HongKong SAR, China +2
Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.
Shijie Lian, Bin Yu, Zhaolong Shen +5
Huazhong University of Science and Technology · Zhongguancun Academy · DeepCybo +5