Discrete image tokenizers provide a sequential interface for vision and multimodal models, but are typically optimized for reconstruction or compression and therefore tend to encode local appearance rather than object-level structure. We introduce COMiT, a communication-inspired framework for learning structured discrete visual representations. COMiT constructs a fixed-length latent message through sequential visual observations: at each step, a transformer processes a localized image crop and updates, refines, and reorganizes the existing token sequence. After several iterations, the resulting message conditions a flow-matching decoder that reconstructs the complete image. The encoder and decoder are implemented within a single transformer and trained end-to-end using flow-matching reconstruction and semantic representation-alignment objectives. COMiT substantially improves compositional generalization and relational reasoning over prior methods. Our experiments show that, while semantic alignment helps ground the representation, attentive sequential tokenization is critical for inducing more interpretable, object-centric token structures.
Figures & tables
Figure 1 : The overall training pipeline of COMiT. A sequence of K random crops is extracted from the input image and iteratively embedded into the latent message mK that is discretized via FSQ Mentzer et al. (2024) . The latter is decoded by the same model using the flow matching objective Lipman et al. (2023) . Additionally, we use REPA Yu et al. (2024b) to speed up the training and SREPA to inject more semantic priors into the latent message.
Table 2
Figure 2 : Effect of attentive tokenization on the visual grounding of tokens. Training with local crops yields much better token–object alignment than training with global crops only. Attention maps from the 10th layer of COMiT-B; both models embed only the global crop.
IN100
MSCOCO
VG
Global crop
# Crops
Ordering
top-1 ↑
top-5 ↑
top-1 ↑
✓
1
–
82.91
41.46
52.11
✓
10
Random
82.30
40.21
52.12
-
9
Random
74.45
38.92
54.96
✓
10
Raster-scan
82.22
41.83
55.78
-
9
Raster-scan
80.62
38.61
51.10
Table 3 : Effect of the cropping policy with COMiT-B. A single global crop gives the best trade-off between performance and test-time cost, while some tasks benefit from additional local crops.
Figure 3 : The difference between the adaptive (top) and global+adaptive (bottom) cropping policies. In both cases COMiT aggregates the crops of the input image (the leftmost column) into the latent message and decodes it to obtain the reconstructed image (the rightmost column, 10 NFE with CFG=7.5 ). The columns in-between show which crops are selected together with immediate single step reconstructions (1 NFE with CFG=1.0 ).
IN1k
IN100
MSCOCO
VG
Method
#params
msg len.
voc.
rFID ↓
PSNR ↑
top-1 ↑
top-5 ↑
top-1 ↑
TiTok-L ( Yu et al., 2024a )
614M
32
212
2.21
15.60
17.26
6.22
26.06
TiTok-B ( Yu et al., 2024a )
172M
64
212
1.70
16.80
19.43
12.64
27.31
TiTok-S ( Yu et al., 2024a )
44M
128
212
1.71
17.52
19.47
6.30
26.81
ALIT ( Duggal et al., 2024 )
229M
256
210
7.39
19.76
29.43
8.56
32.14
FlexTok (d12-d12) ( Bachmann et al., 2025 )
254M
256
64k
4.20
18.41
80.25
38.17
53.75
Table 4 : We evaluate COMiT and several baseline 1D tokenizers on our test suite comprising three representative benchmarks described in Section 4 . COMiT outperforms prior work on semantic probing. † : trained on larger data mixtures.
Figure 4 : The way COMiT adds information to the latent message is inherently compositional.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
COMiT-B
COMiT-L
COMiT-XL
DiT parameters
Depth
12
24
28
Hidden size
768
1024
1152
Heads
12
16
16
MLP ratio
4
4
4
REPA/SREPA parameters
Appendix
Table 5: Architecture and training details of COMiT variants.
Benchmark
Model dim
Depth
Heads
Global Batch Size
ImageNet100
768
2
8
512
MSCOCO
768
2
8
256
Visual Genome
64
2
8
512
Appendix
Table 6: Probe hyperparameters used for all benchmarks.
Figure 5: (a) and (b): Ablation of the reconstruction fidelity under different sampling hyperparameters. (c): Ablation of the bottleneck size with COMiT-B.
Figure 6: Emergence of object-centric tokens in COMiT. We visualize the attention maps of specific tokens in one of the deep layers of the network during the decoding stage. One can see that the tokens naturally attend to the objects and their parts.
Figure 7: Generalization of COMiT to other domains, such as rendered animations or medical images. It can be seen that the model has certain symmetry bias that allows it to reduce the uncertainty about the right hand side of the image when the content on the left hand side is observed. Also notice how the adaptive cropping policy selects the most critical regions for reconstruction.
Figure 8: Visualization of COMiT-XL’s attention maps. The attention maps of the tokens that yield the largest IoUs per sample are shown. With adjusting the thresholding percentage, a remarkable mIoU of 0.58 can be achieved. Note that the model has never seen any segmentation maps or class labels during training.
Training variant
IN100 top-1 ↑
CSSD mIoU ↑
95% interval
Paired difference
Full
82.91
0.527
[0.509,0.546]
–
No SREPA
72.26
0.507
[0.490,0.525]
+0.020[+0.012,+0.028]
No local crops
80.94
0.338
[0.324,0.352]
+0.189[+0.175,+0.204]
Appendix
Table 7 : Three-way ablation of COMiT-B on IN100 (attention probe) and CSSD token localization. All variants report results for a single global crop at test time. The last column is the paired full-minus-variant difference in mIoU with its 95% image-bootstrap interval.
Tokenizer
Message shape
IN100 top-1 ↑
TiTok-S (128 tokens) ( Yu et al., 2024a )
128×12
16.60±0.14
FlexTok (d12-d12) ( Bachmann et al., 2025 )
256×6
72.07±0.22
FlexTok (d18-d18) ( Bachmann et al., 2025 )
256×6
74.32±0.14
FlexTok (d18-d28) ( Bachmann et al., 2025 )
256×6
72.76±0.38
COMiT-B (ours)
256×6
71.66±0.25
COMiT-L (ours)
256×6
76.45±0.28
Appendix
Table 8 : Single affine readout over the flattened frozen message of tokenizers with a flattened message dimensionality of 1,536 (IN100 top-1, mean ± sample SD over five probe seeds; COMiT encodes a single global crop).
Training setting
CSSD mIoU ↑
95% interval
One backpropagated update (default)
0.527
[0.509,0.546]
Up to four backpropagated updates
0.566
[0.545,0.587]
Appendix
Table 9 : CSSD token-localization mIoU of COMiT-B trained with the default stop-gradient heuristic and with gradients through up to four message updates (single global crop at test time; 95% image-bootstrap intervals).
Randomized number of crops
Fixed number of crops
Message after
one update
up to four updates
one update
Global crop
0.0
0.4
0.8
1 local crop
4.7
5.5
36.7
2 local crops
4.7
3.5
9.4
3 local crops
5.1
3.5
3.5
4 local crops
3.5
4.7
3.5
Appendix
Table 10 : Fraction of under-used message positions (%, Hi<8 bits) after the global crop or after the first local crops (raster-scan order unless stated otherwise).
Affine readout (5 seeds)
Attention probe (3 seeds)
Message after
R1
R4
F1
R1
R4
F1
Global crop
56.08 ± 0.28
63.41 ± 0.35
46.78 ± 0.46
74.71 ± 0.54
77.09 ± 0.50
65.40 ± 1.68
1 local crop
26.44 ± 0.43
29.49 ± 0.37
16.90 ± 0.15
27.55 ± 1.50
30.56 ± 0.34
18.53 ± 0.75
2 local crops
41.54 ± 0.42
46.12 ± 0.32
29.64 ± 0.36
46.93 ± 1.15
49.27 ± 0.48
36.54 ± 1.52
4 local crops
52.69 ± 0.36
58.91 ± 0.36
41.97 ± 0.49
59.93 ± 0.69
63.43 ± 0.21
51.81 ± 0.39
9 local crops
64.66 ± 0.31
71.08 ± 0.23
51.90 ± 0.41
72.07 ± 1.55
74.08 ± 0.25
61.79 ± 0.78
Appendix
Table 11 : ImageNet100 top-1 accuracy (%) of probes trained on 200 images per class, mean ± std over probe seeds. R1 and R4: randomized number of crops with one and up to four backpropagated updates; F1: fixed number of four crops with one backpropagated update.
Localized at p=0.5
Mean best
Localized at
Distinct
Norm.
Tokenizer
rate ↑
95% interval
IoU ↑
p=0.25
p=0.75
best tokens
entropy
FlexTok (d18-d18)
21.74
[16.39,27.44]
0.386
79.13
2.17
59 / 256
0.662
COMiT-B, no local crops
17.39
[12.86,22.22]
0.358
74.78
1.30
127 / 256
0.843
COMiT-B, no SREPA
23.48
[18.14,29.17]
0.386
80.87
0.87
117 / 256
0.818
COMiT-B (ours)
25.65
[19.72,31.69]
0.406
83.91
0.43
134 / 256
0.854
COMiT-L (ours)
36.52
[30.13,43.22]
0.433
84.35
3.48
126 / 256
0.836
Appendix
Table 12 : Individual-token localization of the 230 large objects in the COCO subset. An object is localized if the best of the individual binarized token attention maps has IoU>p with its mask. Intervals are 95% bootstrap intervals of the rate at p=0.5 . The last two columns count how many distinct token indices supply the best match over all 230 objects, and the normalized entropy of this distribution. All tokenizers produce 256 tokens per image.
Figure 9: Visualization of the nearest neighbors of a sample from the ImageNet100 validation set in the space of COMiT’s latent messages. One can see that despite unsupervised training and simple probing the learned messages tend to cluster into semantically meaningful groups.
Figure 10: Additional results on the difference between the adaptive (odd rows) and global+adaptive (even rows) cropping policies. The first column depicts the ground truth input image. The last column corresponds to the final reconstruction with 10 NFE. The columns in-between demonstrate how COMiT refines its latent message with incoming information (1 NFE reconstructions are shown).
Discrete visual tokenizers map images to ordered sequences of tokens, providing a natural representation for structural scene descriptions. Most use a fixed sequence length, while adaptive methods often require post-hoc search or choose among a small set of rates that control the length. We propose STROP, a discrete tokenizer that learns both a visual program and its image-dependent active length. A length head is trained with a four-phase curriculum using local rate-distortion probes against frozen DINOv3 features, then predicts the active prefix in a single forward pass. At a matched rate of about 250 nominal bits per crop, the adaptive model improves segmentation over a separately trained fixed-length baseline on four benchmarks (by 1.6-3.1 mIoU), and it also beats a fixed K=32 baseline that uses more bits. STROP programs also yield higher segmentation mIoU than FlexTok, One-D-Piece, and ALIT at similar or higher rates, under the same readout architecture and training protocol. STROP therefore learns useful per-image sequence lengths without post-hoc search or a predefined set of compression rates.
Piotr Wyrwiński, Kacper Dobek, Krzysztof Krawiec
Institute of Computing Science Poznan University of Technology, Poznan, Poland
Unified visual tokenization faces a fundamental trade-off between high-fidelity pixel reconstruction (spatial equivariance) and semantic abstraction (conceptual invariance). We attribute this conflict to Manifold Misalignment: naive joint optimization induces opposing gradients, creating a zero-sum game between reconstruction and perception. To address this, we propose MUSE, a framework based on Topological Orthogonality. By treating Structure as an orthogonal bridge, MUSE decouples optimization within Transformers: structural gradients refine attention topology, while semantic gradients update feature values. This turns destructive interference into Mutual Reinforcement. Experiments show that MUSE breaks the trade-off, achieving state-of-the-art generation quality (gFID 3.08) and surpassing its teacher InternViT-300M in linear probing (85.2% vs. 82.5%), demonstrating that structurally aligned reconstruction can enhance semantic perception. Code is available at https://github.com/PanqiYang1/MUSE.
Panqi Yang, Haodong Jing, Jiahao Chao +5
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence,National Engineering Research Center of Visual Information and Applications,and Institute of Artificial Intelligence and Robotics, Xi’an Jiao Tong University · Xiaohongshu Inc.
Leading flexible vision tokenizers achieve SOTA quality at an extreme cost, relying on parameter-heavy backbones and slow, multi-step generative decoders. We depart from this complex, spatial-token paradigm and introduce a simple, lightweight, and fast channel-wise flexible-length tokenizer. Our method treats each latent channel as a visual token, enabling a parameter-efficient CNN-Transformer hybrid backbone. Furthermore, employing a stochastic tail-dropping paradigm during training naturally forces channels to organize by semantic importance. This allows for flexible compression at inference by simply retaining the first k channels, and naturally enables variable-length autoregressive image generation. We validate our approach through extensive experiments on ImageNet, demonstrating consistent quality across diverse token budgets. The results establish a new quality-efficiency frontier: our model achieves state-of-the-art perceptual quality (rFID 2.92) while being 8.6× faster in decoding and 2.1× smaller (159M params) than the next-best alternative. Our work establishes channel-wise tokenization as a powerful and practical paradigm for efficient visual representation. Project page: https://channeltok.github.io