Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action tokenizer that reframes tokenization as a causally structured generative process. CATok introduces a conditional annealing mechanism that extracts action tokens by progressively annealing a flow-matching process: each token is conditioned on all preceding tokens and encodes the residual reconstruction signal at a specific noise level, establishing a coarse-to-fine causal token space whose generative semantics are structurally aligned with autoregressive modeling. A token-conditioned flow-matching decoder built on Multimodal Diffusion Transformer (MMDiT) reconstructs continuous action chunks from these discrete tokens with the precision of hybrid diffusion-head architectures. This discrete bottleneck enforces knowledge insulation by design, cleanly separating high-level semantic reasoning from low-level motor execution without requiring explicit attention masking. Extensive evaluations across three simulation benchmarks and real-world robotic manipulation tasks demonstrate that CATok consistently surpasses existing tokenization methods in both reconstruction fidelity-compression tradeoff and inference efficiency, while improving VLA task success rate and training efficiency, establishing a high-performance, scalable foundation for purely autoregressive VLA systems.
Figures & tables
Figure 1 : Flow-Matching to Action Tokens. CATok grounds discrete action tokens in the continuous flow matching trajectory. Each token is learned as a stage-wise information increment, imbuing the sequence with a natural coarse-to-fine hierarchy. By construction, these increments follow the temporal evolution of the flow, ensuring a causally-ordered structure that seamlessly aligns with autoregressive generation.
Figure 2 : An Overview of CATok . (a) CATok Pipeline. CATok encodes an action chunk with a Dual-Stream Encoder into latent queries, discretizes them via a Bottleneck VQ , and reconstructs actions using a Flow-Matching Decoder with Conditional Annealing . (b) Conditional Annealing Mechanism progressively masks the first κ(t) token embeddings according to flow matching timestep t: darker colors indicate tokens earlier in the sequence, and lighter colors indicate later tokens, assigning each token to a different stage of the flow trajectory.
Figure 3 : An Overview of VLA- CATok . VLA- CATok feeds visual, language, and proprioceptive tokens into a VLM backbone to autoregressively predict causal action tokens. These tokens are mapped via the learned codebook and processed by a frozen MMDiT flow-matching decoder to produce the final action in a single denoising step. Discrete tokens serve as the sole interface, naturally insulating the VLM’s pretrained knowledge from action-specific gradients.
Figure 4 : Real-world scene layouts for the three task suites. From left to right: Pick-Spatial , testing spatial generalization by varying cup and plate positions; Pick-Color , testing color generalization with different unseen cup colors; and Stack-Long , testing long-horizon manipulation under different spatial layouts.
LIBERO
SimplerEnv
RoboTwin 2.0
Method
Spatial
Object
Goal
Long
Avg.
Spoon
Carrot
Stack
Eggplant
Avg.
Clean
Randomized
BIN
0.586
0.878
0.680
0.604
0.687
0.542
0.333
0.208
0.708
0.448
0.213
0.221
FAST
0.960
0.998
0.962
0.901
0.955
0.500
0.375
0.375
0.625
0.469
0.478
0.478
OAT
0.428
0.876
0.704
0.276
0.571
0.417
0.208
0.167
0.667
0.365
0.229
0.233
CATok
0.978
0.994
0.954
0.910
0.959
0.458
0.417
0.542
0.542
0.490
0.489
0.531
Table 1 : Comparison on robotic manipulation simulation benchmarks. CATok achieves the best overall performance across all benchmarks, with particularly significant improvements on long-horizon tasks such as Long in LIBERO and StackGreenCubeOnYellowCube in SimplerEnv, demonstrating the effectiveness of causal action tokenization for modeling long-range temporal dependencies.
Figure 5 : (a) Training efficiency on the SimplerEnv benchmark. Curves are Gaussian-smoothed to reduce evaluation stochasticity without changing the relative ranking of different methods. Notably, CATok reaches the best baseline performance using only around 50% of the training steps. (b) VLA inference latency. By jointly considering autoregressive generation and decoding overhead, CATok achieves near-minimal end-to-end inference latency while attaining the best success rate among all methods.
Method
Pick-Spatial
Pick-Color
Stack-Long
FAST
0.450
0.300
0.275
CATok
0.650
0.500
0.475
Table 2 : Real-world manipulation results. CATok consistently outperforms FAST across spatial, color, and long-horizon generalization tasks.
Figure 6 : CATok Produces Causal Tokens with Generative Semantics . (a) Causal Ordering CATok ’s entropy decreases in the forward order ( CATOK fwd ) and increases in the reverse order ( CATOK rev ), while FAST and Bin show weaker positional structure. (b) Token embedding geometry. T-SNE map shows slot-dependent organization, progressing from smooth early-token regions to compact late-token clusters. (c) Prefix reconstruction. Increasing token prefixes produce structured trajectory updates, with dashed curves showing incremental contributions.
Figure 7 : Intrinsic properties of action tokenizers. (a–c) Fidelity vs. compression efficiency. VRR measures reconstruction quality; CR measures compression ratio; VRR × CR jointly captures the fidelity–compression tradeoff. CATok 8 achieves the highest VRR × CR ( 16.85 ), outperforming all baselines at the same compression level. (d–e) Inference efficiency. Encode and decode latencies are shown on a log scale. The subscript k in OAT k and CATok k denotes the number of tokens.
Table 4: Controlled ablation of conditional annealing, random masking, and decoder architecture.
Figure 8 : Effect of the number of action tokens on downstream VLA performance on the LIBERO benchmark. Performance improves from 8 to 16 tokens but degrades with larger token budgets, with 16 tokens achieving the best overall performance.
Figure 9 : Causality validation via native-decoder interventions. Both donor-swap (a) and single-token removal (b) exhibit a clear position-dependent structure: early tokens have stronger influence on coarse action components and earlier generation stages, whereas later tokens increasingly affect fine-grained action details, confirming that the learned token ordering corresponds to meaningful causal roles rather than an arbitrary positional ordering.
The Hong Kong Polytechnic University, HongKong SAR, China · Shanghai AI Laboratory, Shanghai, China · The University of Hong Kong, HongKong SAR, China +2