cs.ROOct 6, 2026

CAP: Codebook-Aligned Prediction for Tokenized Robot Policies

Authors: Haoran Chen, Jingtian Ji, Samuel Wheeler, Kaylene Caswell Stocking, Matthew Walter

Organizations: Toyota Technological Institute at Chicago · Argonne National Laboratory

Abstract

Action tokenization converts continuous robot actions into discrete symbols that can be modeled autoregressively. However, existing tokenizer-based policies typically ignore the tokenizer's learned latent code structure: after tokenization, the policy treats tokens as unrelated class indices and learns a new classifier from scratch. We show that this discarded structure is valuable. We introduce Codebook-Aligned Prediction (CAP), a method that directly reuses the tokenizer's code vectors as policy class prototypes while leaving the tokenizer and policy backbone otherwise unchanged. Across four quantizer families, three simulation benchmarks, and two real-robot tasks, CAP consistently improves task success over standard token classification heads while holding the tokenizer (and therefore its reconstruction quality) fixed. Our analysis further shows that these gains are not explained by higher token accuracy or changes in the policy head alone. Instead, reusing the tokenizer codebook provides the policy with valuable information about the tokenizer's learned latent structure across tokens, making token prediction errors more benign in action space and improving the representations learned by the policy backbone. These results suggest that action tokenizers learn useful action-aware latent structure beyond discrete targets that should be preserved when training downstream policies.

Explore similar work

CardsList
  1. Beyond Reconstruction: What Matters in Action Tokenization for Robot Policies?

    Oct 6, 2026Haoran Chen, Jingtian Ji, Samuel Wheeler +2Discrete Action TokenizersRobot Policies

  2. Ordered Action Tokens for Visuomotor Policy Learning

    Jul 23, 2026Chaoqi Liu, Yue Zhao, Haonan Chen +4Discrete Action TokenizersVisuomotor Policy