cs.ROSep 28, 2026

Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching

Authors: Chenyu Zhang, Yuhang Cao, Daru Du, Yingxi Lu, Jing Shao, Ruoqu Chen, Jiajun Liu, Liu Cao, +3 more

Organizations: IIIS, Tsinghua University · Shanghai Qizhi Institute

Abstract

Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action tokenizer that reframes tokenization as a causally structured generative process. CATok introduces a conditional annealing mechanism that extracts action tokens by progressively annealing a flow-matching process: each token is conditioned on all preceding tokens and encodes the residual reconstruction signal at a specific noise level, establishing a coarse-to-fine causal token space whose generative semantics are structurally aligned with autoregressive modeling. A token-conditioned flow-matching decoder built on Multimodal Diffusion Transformer (MMDiT) reconstructs continuous action chunks from these discrete tokens with the precision of hybrid diffusion-head architectures. This discrete bottleneck enforces knowledge insulation by design, cleanly separating high-level semantic reasoning from low-level motor execution without requiring explicit attention masking. Extensive evaluations across three simulation benchmarks and real-world robotic manipulation tasks demonstrate that CATok consistently surpasses existing tokenization methods in both reconstruction fidelity-compression tradeoff and inference efficiency, while improving VLA task success rate and training efficiency, establishing a high-performance, scalable foundation for purely autoregressive VLA systems.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SA-VLA: State-aware tokenizer for improving Vision-Language-Action Models' performance

    Jun 29, 2026Tengyue Jiang, Chunpu Xu, Jiayue Kang +1Discrete Action TokenizersGeneralizable Vision-Language-Action Policies

  2. M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models

    Sep 16, 2026Chunpu Xu, Zhixuan Liang, Yuhao Zhang +6Discrete Action TokenizersDiffusion-Based Vision-Language-Actions

  3. ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

    Sep 16, 2026Shijie Lian, Bin Yu, Zhaolong Shen +5Discrete Action TokenizersDiffusion-Based Vision-Language-Actions