Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly that: the first coarse tokens carry the clip's global semantics while later tokens further specify details. Existing flexible tokenizers only apply a representation-alignment (REPA) loss on early decoder hidden states, a target the decoder can partly meet from its noised input instead. We introduce SemanTok, a flexible video tokenizer that feeds frozen DINO features into its encoder and adds lightweight heads that reconstruct them from each retained token prefix alone. SemanTok achieves high semantic alignment and video fidelity at every AR model size: a 201M SemanTok AR model matches or beats a VideoFlexTok AR model 3.4× its size, and larger SemanTok AR models further improve fidelity. It keeps semantic alignment on out-of-distribution classes and gives the decoder higher semantic alignment at every noise level, including pure noise. It performs well in both reconstruction and generation, and its short token prefixes are cheaper to predict and give better generation fidelity, with pixel detail deferred to later tokens.
Figures & tables
Figure 1: From a prompt with “…orange basketball…” (left, uCO3D), our SemanTok tokenizer lets a 201M autoregressive (AR) model generate a video that keeps the ball’s shape and appearance through the orbit with a budget of only k=4 tokens per latent frame. VideoFlexTok: at the same model’s size the ball is semantically misaligned, and even an 11× larger AR model (2.29B) leaves its shape unstable up to k=64 . So the smaller, faster model on SemanTok tokens matches or beats the larger one of VideoFlexTok. Right (Kinetics-600, class “yoga”): both tokenizers at the same 2.29B AR size, i.e., equal cost. SemanTok keeps a complex body motion stable from k=16 , and larger k refines its appearance and motion. Meanwhile VideoFlexTok changes the scene between k=4 and k=16 and poorly simulates body motion. Both VideoFlexTok and SemanTok are variable-length: the AR model generates tokens coarse to fine, and any budget k decodes to a video. SemanTok prioritizes semantic content in its first tokens, so a small k already fixes what the clip contains. Pipeline in fig. 2 ; videos on the project page.
Figure 2: Overview. (1) Tokenizer training: encoder, FSQ, and diffusion decoder are trained jointly, with nested dropout so that the decoder can reconstruct the clip from any token prefix. SemanTok keeps this VideoFlexTok recipe ( orange ) and adds semantic supervision from a frozen teacher: its features enter the encoder, and every retained prefix is trained to predict them ( purple ; details in fig. 3 ). (2) AR training and generation: with the tokenizer frozen, an AR model learns to predict its tokens in time-first, coarse-to-fine order, and the decoder trained in (1) renders any generated prefix as video.
Figure 3: The figure contrasts two tokenizer variants. VideoFlexTok provides the base coarse-to-fine path ( orange ), while SemanTok changes its encoder input and adds a Dense DINO head and a Class DINO head ( purple ). A frozen VidTok VAE maps the clip to latents; each frame projects its patches to e and packs them with K learnable register tokens r : (et,0,…,et,P−1,rt,0,…,rt,K−1) , interleaved over time. The time-causal encoder transforms the packed sequence, retaining only the register-token outputs for FSQ; nested dropout forms the kept token prefix that conditions a time-causal decoder reconstructing VAE latents under flow-matching and decoder-REPA losses. SemanTok concatenates frozen DINOv2 patch features with each VAE patch before projection, and adds a zero-initialized projection of the matching DINO class token to rt,0 . Readout queries in the Dense DINO head and Class DINO head cross-attend to the kept token prefix from frames t′≤t .
Figure 4: 1 SemanTok improves generation at every compute budget on Kinetics-600 and uCO3D. Generation versus AR inference FLOPs per clip. Each faded curve is one AR size (colour) sweeping the token budget k from 1 to 256 tokens per frame. Black: the best score each tokenizer reaches at a given compute. SemanTok’s envelope is better over most of the compute range.
Figure 5
Figure 7: 2 SemanTok stays ahead at both AR model sizes.Class-to-video rollouts for one Kinetics-600 playing-guitar label. At 201M, SemanTok degrades slowly over time, while VideoFlexTok is sharp only at t=1 . Videos on the project page.
Figure 8: 3 SemanTok reconstructions on uCO3D from ground-truth tokens recover the original object class at an earlier k than VideoFlexTok; for both in-distribution (ID) clips and out-of-distribution (OOD) clips.
Figure 9: 4 SemanTok’s decoder-REPA readout depends less on the noised latent and achieves higher REPA-alignment than VideoFlexTok’s.
Figure 9
Figure 11: 6 SemanTok’s generation gain comes specifically from its first tokens, which are cheaper to predict (Kinetics-600). (a) Δ gFVD after teacher-forcing the first m ground-truth tokens per frame and free-running to k=256 : at 201M, SemanTok leads by 23% with nothing forced, and forcing only the first 16–64 tokens removes the lead. (b) Under a 201M AR model, SemanTok’s cross-entropy per token position is lower for the first 128 positions and higher over 129–192. (c) Larger AR models narrow the gap between generation and reconstruction for both tokenizers, but none closes it.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 12: Tokenizer reconstruction on uCO3D from ground-truth tokens: two in-distribution clips and one out-of-distribution clip. PSNR and ClipV against VAE-GT; bold is best per k . Videos: ReconstructionUCO3D.mp4 , ReconstructionUCO3D_Flashlight.mp4 .
Figure 13: Text-to-video on uCO3D, 201M AR model; one rollout per tokenizer, decoded from its first k tokens per frame. Prompts: “A red fire extinguisher with a black handle and silver nozzle, sits on a countertop. It has a white label on its side and a black strap around its middle.” and “A black fedora hat with a short brim and small, round crown sits on a white surface. It has two small holes at the top of the crown.” Video: GenerationUCO3D_FireExtinguisher.mp4 .
Figure 14: Tokenizer reconstruction on Kinetics-600 from ground-truth tokens. PSNR against VAE-GT; bold is best per k . Videos: ReconstructionK600_HockeyStop.mp4 , ReconstructionK600_Luge.mp4 .
Figure 15: Class-to-video on Kinetics-600 (playing guitar) with a 49M AR model; one rollout per tokenizer (the first evaluation sample of the class), decoded from its first k tokens per frame. Video: GenerationK600_Guitar_49M.mp4 .
Figure 16: Class-to-video on Kinetics-600 (cooking egg) across AR model sizes at k=16,32,64 , frame t=9 . Each (size, tokenizer) is one rollout decoded from its first k tokens per frame. SemanTok’s is the first evaluation sample of the class; VideoFlexTok’s is, per size, the one of its rollouts of the class closest to SemanTok’s in ViCLIP embedding. Video: GenerationK600_CookingEgg_k16-64_ARsize.mp4 .
Figure 17: Text-to-video on uCO3D (out-of-distribution category) across AR model sizes at k=8,32,256 , frame t=9 . Each (size, tokenizer) is that model’s evaluation rollout for the same prompt, decoded from its first k tokens per frame. Prompt: “A light blue dumbbell with a hexagonal handle and a rounded head sits on a surface. It has white text or numbers printed on its side, such as ‘Tone’ or ‘5 LB’. The overall color scheme is dominated by the blue of the dumbbell itself.” Video: GenerationUCO3D_Barbell_ARsize_x_k.mp4 .
1
4
8
16
32
64
128
256
Kinetics-600
rFVD↓
VideoFlexTok
426.7
262.0
203.0
148.1
125.2
105.5
79.1
64.8
SemanTok
427.8
217.4
178.3
156.0
143.9
129.3
99.6
68.2
ViCLIP ↑
VideoFlexTok
0.155
0.163
0.168
0.174
0.177
0.182
0.185
0.188
SemanTok
0.170
0.179
0.182
0.185
0.186
0.187
0.190
0.192
ClipV ↑
VideoFlexTok
0.569
0.631
0.678
0.730
0.767
0.804
0.840
0.867
Appendix
Table 2: Tokenizer reconstruction versus k . Bold is better.
Figure 18: Where the tokenizers keep DINO semantics. Decoder-REPA readout of DINOv2 from the first k tokens per frame, 256 validation clips per dataset. Rows: Class-to-Video (Kinetics-600) and Text-to-Video (uCO3D). Colour: decoder noise level σ ; at σ=1 the input is pure noise, so the prefix is the decoder’s only source. (a, d) VideoFlexTok. (b, e) SemanTok; dotted: its dense token head, which reads the prefix without the decoder. (c, f) SemanTok minus VideoFlexTok at each σ ; positive favors SemanTok. Figure 9 overlays (a, b) and (d, e) and adds the gain from σ=1 to σ=0.25 .
Knob
uCO3D
Kinetics-600
Tokenizer training tokens
66B
131B
RGB frames / latent frames
17 / 5
17 / 5
Latent grid / channels
16×16 / 16
same
Encoder / decoder width
1152 / 1152
same
Tokens per latent frame
256
256
FSQ levels / vocabulary
[8,8,8,5,5,5] / 64k
same
Appendix
Table 3: Tokenizer settings used by the reported checkpoints.
Figure 19: SemanTok’s lead holds throughout tokenizer training on both datasets. Each point is an AR model trained on one tokenizer checkpoint’s tokens. Top: uCO3D, a 49M AR probe trained for 8k steps, official validation split. Bottom: Kinetics-600, a 201M AR model with the recipe and evaluation of fig. 4 . No gap closes as the tokenizer trains longer.
Non-emb. params
Depth
Width
Heads
uCO3D peak LR
K600 peak LR
49M
10
640
10
8.00×10−4
1.60×10−3
85M
12
768
12
6.67×10−4
1.33×10−3
201M
16
1024
16
5.00×10−4
1.00×10−3
393M
20
1280
20
4.00×10−4
8.00×10−4
679M
24
1536
24
3.33×10−4
6.67×10−4
1.33B
30
1920
30
2.67×10−4
5.33×10−4
Appendix
Table 4: AR model size ladder. uCO3D uses 0.512/width ; Kinetics-600 uses the VideoFlexTok rule 1.024/width .
Knob
Dataset-specific value
Conditioning
uCO3D: umT5, 128 tokens; K600: class
Tokenizer crop views
uCO3D: 25; K600: 6
AR training tokens
uCO3D: 13.1B at every size
K600: 13.1B through 679M; 26.2B thereafter; 65.5B for fig. 6
Trunk dropout
uCO3D: 0.25; K600: 0.1
Condition drop
uCO3D: 0.2; K600: 0.1
Appendix
Table 5: Dataset-specific AR settings. Crop views affect only AR training; validation tokens are unaugmented.
Figure 20: Token repeats by position on Kinetics-600 (running mean over 16 positions). Solid: duplicates of an earlier token in the same frame. Dotted: copies of the previous frame’s token at the same position.
Figure 21: SemanTok’s advantage over VideoFlexTok at equal AR size, as a function of the token budget. Each panel shows the paired difference between the two tokenizers under the same AR size, k , prompts, and labels, signed so that positive values favour SemanTok (VideoFlexTok minus SemanTok for gFVD and gFID; SemanTok minus VideoFlexTok for class accuracy and ClipV). Lines are point estimates for three AR sizes; bands are 95% bootstrap intervals over evaluation clips; the grey line marks no difference. SemanTok’s semantic-alignment lead appears from k=2 , its fidelity lead from k≈16 on uCO3D and from k=4 (small AR models) to k=32 (largest) on Kinetics-600, and VideoFlexTok is ahead only at k≤4 on uCO3D. Diamonds: Kinetics-600 at k=16 with six sampling seeds pooled.
Figure 22: SemanTok’s advantage for every AR size at five token budgets. Rows fix k , columns fix the metric, and within each panel the AR size grows from 49M (left) to 2.29B (right). Dots are the paired difference at equal AR size and k , signed as in fig. 21 so that positive values favour SemanTok; whiskers are 95% bootstrap intervals over evaluation clips. Teal marks an interval above zero, grey an interval that includes zero, and magenta an interval below zero. Each column shares its y -range, so the change with k reads down the page. The Kinetics-600 k=16 row pools six sampling seeds; all other cells use one.
AR
k=1
4
8
16
32
64
128
256
gFVD ↓
49M
VideoFlexTok
390.8
295.7
317.5
303.9
338.8
478.6
588.9
653.7
SemanTok
390.4
236.8†
239.2
223.7†
223.7†
324.0
450.2
610.1†
85M
VideoFlexTok
386.3
284.1
293.3
278.9
307.0
450.6
548.9
627.3
SemanTok
385.4
244.0†
228.8
219.9†
212.5†
322.8
411.0
554.5†
201M
VideoFlexTok
384.4
278.9
279.9
272.9
302.9
420.7
501.0
560.3
Appendix
Table 6: Kinetics-600 class-to-video generation: fidelity. AR models sampled with k tokens per frame, tokenizers at 200k steps, AR models at 20k steps (1.33B and 2.29B at 40k), n=2048 . Bold is better at that k and size; † marks the better tokenizer where the 95% paired-bootstrap CI of the difference (single sampling seed) excludes zero. CIs exist at k∈{1,4,16,32,256} only. ‡ marks SemanTok gains that exclude zero once six sampling seeds are pooled; with pooling, every k=16 gFVD and gFID gain of SemanTok excludes zero.
AR
k=1
4
8
16
32
64
128
256
ViCLIP ↑
49M
VideoFlexTok
0.179
0.192
0.187
0.189
0.187
0.161
0.160
0.156
SemanTok
0.193
0.208
0.206
0.208
0.206
0.179
0.175
0.168
85M
VideoFlexTok
0.178
0.193
0.188
0.191
0.190
0.164
0.161
0.159
SemanTok
0.194
0.208
0.208
0.208
0.206
0.181
0.177
0.171
201M
VideoFlexTok
0.178
0.194
0.191
0.194
0.192
0.165
0.163
0.161
Appendix
Table 7: Kinetics-600 class-to-video generation: semantic alignment. Same samples as Table 6 . Class acc. is UMT-L top-1. Bold is better at that k and size; † marks the better tokenizer where the 95% paired-bootstrap CI of the difference (single sampling seed) excludes zero. CIs exist for class acc. at k∈{1,4,16,32,256} only; ViCLIP and ClipV have none.
AR
k=1
4
8
16
32
64
128
256
gFVD ↓
49M
VideoFlexTok
436.3
259.9
235.2
232.9
254.7
284.5
344.1
379.1
SemanTok
452.0
258.7
226.9
226.2
218.2†
230.4†
281.2†
346.9†
85M
VideoFlexTok
426.8†
255.9
229.0
237.1
260.1
275.5
338.4
367.1
SemanTok
461.9
239.4
233.9
203.7†
209.6†
218.9†
261.1†
320.1†
201M
VideoFlexTok
424.4
244.2
211.8
218.6
224.5
244.0
286.8
315.5
Appendix
Table 8: uCO3D text-to-video generation: fidelity. Tokenizers at 100k steps, AR models at 20k steps; 1,014 ID and 152 OOD clips pooled, gFVD and gFID as one Fréchet distance. Bold is better at that k and size; † marks the better tokenizer where the 95% paired-bootstrap CI of the difference (single sampling seed) excludes zero. CIs cover every cell.
AR
k=1
4
8
16
32
64
128
256
ViCLIP ↑
49M
VideoFlexTok
0.170
0.210
0.217
0.221
0.221
0.220
0.217
0.212
SemanTok
0.169
0.205
0.214
0.226
0.230
0.231
0.226
0.220
85M
VideoFlexTok
0.170
0.208
0.218
0.221
0.222
0.222
0.219
0.216
SemanTok
0.170
0.206
0.215
0.227
0.230
0.232
0.228
0.222
201M
VideoFlexTok
0.169
0.208
0.217
0.222
0.223
0.224
0.221
0.220
Appendix
Table 9: uCO3D text-to-video generation: semantic alignment. Same samples as Table 8 . Class acc. is NCM. Bold is better at that k and size; † marks the better tokenizer where the 95% paired-bootstrap CI of the difference (single sampling seed) excludes zero. CIs cover every ClipV and class-acc. cell; ViCLIP has none.
σ
k=1
4
8
16
32
64
128
256
Kinetics-600
1
VideoFlexTok
0.480
0.529
0.555
0.581
0.599
0.614
0.624
0.631
SemanTok
0.491
0.565
0.593
0.623
0.646
0.668
0.688
0.707
0.75
VideoFlexTok
0.517
0.552
0.571
0.593
0.608
0.621
0.630
0.636
SemanTok
0.527
0.584
0.608
0.634
0.654
0.673
0.692
0.709
0.5
VideoFlexTok
0.612
0.624
0.630
0.639
0.645
0.650
0.654
0.656
Appendix
Table 10: Decoder-REPA readout versus k and decoder noise σ . Cosine between DINOv2 features of the first frame of each latent group and the decoder-REPA readout from the first k tokens per frame, 256 validation clips. σ=1 is pure noise, as in generation. Bold is better at that k and σ .