cs.CVFeb 24, 2026

COMiT: Learning Structured Visual Tokens through Sequential Communication

Authors: Aram Davtyan, Yusuf Sahin, Yasaman Haghighi, Sebastian Stapf, Pablo Acuaviva, Alexandre Alahi, Paolo Favaro

Organizations: Computer Vision Group, University of Bern, Switzerland · VITA Lab, EPFL, Switzerland

Abstract

Discrete image tokenizers provide a sequential interface for vision and multimodal models, but are typically optimized for reconstruction or compression and therefore tend to encode local appearance rather than object-level structure. We introduce COMiT, a communication-inspired framework for learning structured discrete visual representations. COMiT constructs a fixed-length latent message through sequential visual observations: at each step, a transformer processes a localized image crop and updates, refines, and reorganizes the existing token sequence. After several iterations, the resulting message conditions a flow-matching decoder that reconstructs the complete image. The encoder and decoder are implemented within a single transformer and trained end-to-end using flow-matching reconstruction and semantic representation-alignment objectives. COMiT substantially improves compositional generalization and relational reasoning over prior methods. Our experiments show that, while semantic alignment helps ground the representation, attentive sequential tokenization is critical for inducing more interpretable, object-centric token structures.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Structure over Pixels: Learning Variable-Length Visual Programs

    May 26, 2026Piotr Wyrwiński, Kacper Dobek, Krzysztof KrawiecVisual TokenizersSemantic Scene Understanding

  2. MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality

    May 7, 2026Panqi Yang, Haodong Jing, Jiahao Chao +5Visual TokenizersSelf-Supervised Vision Transformers

  3. ChannelTok: Efficient Flexible-Length Vision Tokenization

    Jun 3, 2026Sukriti Paul, Arpit Bansal, Tom GoldsteinVisual TokenizersCnn-Transformer Tradeoff