cs.CVSep 30, 2026

SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

Authors: Mikhail Dereviannykh, Vikram Voleti, Simon Donne, Mallikarjun Byrasandra Ramalinga Reddy, Shimon Vainer, Mark Boss

Organizations: Stability AI · Karlsruhe Institut für Technologie

Abstract

Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly that: the first coarse tokens carry the clip's global semantics while later tokens further specify details. Existing flexible tokenizers only apply a representation-alignment (REPA) loss on early decoder hidden states, a target the decoder can partly meet from its noised input instead. We introduce SemanTok, a flexible video tokenizer that feeds frozen DINO features into its encoder and adds lightweight heads that reconstruct them from each retained token prefix alone. SemanTok achieves high semantic alignment and video fidelity at every AR model size: a 201M SemanTok AR model matches or beats a VideoFlexTok AR model 3.4×3.4\times its size, and larger SemanTok AR models further improve fidelity. It keeps semantic alignment on out-of-distribution classes and gives the decoder higher semantic alignment at every noise level, including pure noise. It performs well in both reconstruction and generation, and its short token prefixes are cheaper to predict and give better generation fidelity, with pixel detail deferred to later tokens.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging

    May 29, 2026Luyuan Zhang, Siyuan Li, Zedong Wang +7Visual TokenizersVisual Generation

  2. End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer

    May 1, 2026Wenda Chu, Bingliang Zhang, Jiaqi Han +4Autoregressive Image GenerationVisual Tokenizers

  3. VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations

    Apr 27, 2026Maitreya Patel, Jingtao Li, Weiming Zhuang +2Autoregressive Image GenerationVisual Tokenizers