Scaling Video Generation for Reasoning: At What Cost?
Organizations: Rice University · Carnegie Mellon University
Abstract
We study whether scaling video generation enables models to reason about hidden information from the past frames, and at what computational cost. Our controlled benchmark requires predicting nine prescribed moves of an initially solved 2x2x2 Rubik's Cube from a fixed view of three faces. Correct predictions require inferring how actions change hidden states, and the simulator provides exact ground truth for evaluation. Models learn plausible cube geometry early, while correct sticker configurations require substantially more training. Although validation MSE follows approximate power-law scaling, lower MSE loss does not reliably indicate downstream reasoning capabilities. Smaller autoregressive models achieve higher state accuracy with limited compute, while larger models reach higher accuracy after more training. At roughly 0.1 PF-days, the 70M-parameter model correctly predicts the visible sticker configuration in 44.6% of post-action frames, compared with 0.3% for the 1B model, which reaches 83.7% at 3.14 PF-days. Symbolic state supervision raises the 20M model's frame accuracy from 31.1% to 67.3% at the same training-data budget, suggesting that learning representations of state changes can complement scaling.
Figures & tables
| Metric | Definition |
|---|---|
| MSE loss | Mean squared error between the predicted and target flow velocities in video latent space on validation set. |
| Action following acc. | Fraction of actions in the prompt that are correctly applied to the cube. |
| Sticker acc. | Fraction of visible stickers ( stickers visible faces) with the correct color after each action. |
| Frame acc. | Fraction of the nine post-action frames in which all 12 visible stickers have the correct color. |
| Full-trajectory acc. | Fraction of episodes in which all 12 visible stickers are correct after every one of the nine actions. |
| Configuration | Layers | Heads | Parameters (M) | Generation | Action | Sticker | Frame | ||
|---|---|---|---|---|---|---|---|---|---|
| Deep-narrow | 512 | 22 | 8 | 1408 | 95.03 | AR- | 17.22 | 25.37 | 3.67 |
| AR- | 6.33 | 22.20 | 0.78 | ||||||
| Bidir | 7.33 | 21.59 | 0.33 | ||||||
| Balanced | 640 | 14 | 10 | 1728 | 94.20 | AR- | 72.44 | 52.00 | 24.33 |
| AR- | 9.67 | 22.46 | 1.67 | ||||||
| Bidir | 6.78 | 21.55 | 0.56 |
| Generation | PF-days (approx.) | Frame accuracy after 20 actions (%) | Full-trajectory accuracy (%) |
|---|---|---|---|
| AR | 94.01 | 92.81 | |
| Bidir | 99.86 | 99.85 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Video | Sticker | Frame | Full trajectory |
|---|---|---|---|
| Original RGB | 100.00 | 100.00 | 100.00 |
| Wan VAE reconstruction | 100.00 | 100.00 | 100.00 |
| Source | Action following (all) | Action following (A9) |
|---|---|---|
| VAE reconstruction | 92.6 | 92.1 |
| AR | 80.1 | 80.6 |
| Bidirectional | 69.6 | 67.7 |
| Component | Value or rule | Scaling treatment |
|---|---|---|
| Transformer block | Self-attention, cross-attention, SwiGLU | Fixed |
| Model width | Tables 2 and 7 | Scaled |
| Depth | Selected wide–shallow ladder | Scaled |
| Position encoding | 3D RoPE | Fixed |
| Head width | Fixed | |
| Attention heads | Derived |
| Model name | Parameters (M) | Layers | Heads | ||
|---|---|---|---|---|---|
| 20M | 19.70 | 384 | 8 | 6 | 1024 |
| 30M | 33.45 | 448 | 10 | 7 | 1216 |
| 70M | 67.81 | 640 | 10 | 10 | 1728 |
| 120M | 115.8 | 768 | 12 | 12 | 2048 |
| 270M | 271.8 | 1088 | 14 | 17 | 2944 |
| 550M | 548.1 | 1408 | 17 | 22 | 3776 |
| Generation | Peak learning rate | ||
|---|---|---|---|
| AR- | 0.1678 | 0.1217 | 0.0768 |
| AR- | 0.1790 | 0.1337 | 0.0810 |
| AR- | 0.1715 | 0.1287 | 0.0783 |
| Bidir | 0.1789 | 0.1274 | 0.0781 |
| Shape | AR- | AR- | Bidir |
|---|---|---|---|
| Deep–narrow | 5.015 | 5.281 | 6.343 |
| Balanced | 4.542 | 4.753 | 5.527 |
| Wide–shallow | 4.367 | 4.548 | 5.150 |
| Shape | Family | MSE ( ) | Action | Sticker | Frame |
|---|---|---|---|---|---|
| deep–narrow | AR-K1 | 2.233 | 17.22 | 25.37 | 3.67 |
| deep–narrow | AR-K4 | 2.562 | 6.33 | 22.20 | 0.78 |
| deep–narrow | BIDIR | 3.212 | 7.33 | 21.59 | 0.33 |
| balanced | AR-K1 | 1.883 | 72.44 | 52.00 | 24.33 |
| balanced | AR-K4 | 2.533 | 9.67 | 22.46 | 1.67 |
| balanced | BIDIR | 3.426 | 6.78 | 21.55 | 0.56 |
| Model | Mean frame: Free | GT prefix | A9 frame: Free | GT prefix |
|---|---|---|---|---|
| XS | 33.89 | 38.67 | 0.00 | 5.00 |
| S | 42.22 | 47.89 | 0.00 | 9.00 |
| M | 52.56 | 60.00 | 0.00 | 22.00 |
| B | 48.11 | 55.00 | 1.00 | 9.00 |
| L | 63.56 | 71.67 | 3.00 | 38.00 |
| XL | 64.56 | 73.33 | 5.00 | 45.00 |
| Operation | Parameters | Forward FLOPs per target token |
|---|---|---|
| Language projection | ||
| Latent input/output projections | ||
| Target visual Q/K/V and output | ||
| Visual-context K/V (shared weights) | — | |
| Visual attention products | — | |
| Language cross-attention projections |
| Metric | Segment | |||
|---|---|---|---|---|
| AR frame | Middle third | 0.336 | 0.2178 | |
| AR frame | Final third | 0.5607 | 1.055 | |
| AR trajectory | Middle third | 0.9501 | 0.1185 | |
| AR trajectory | Final third | 1.59 | 0.7931 | |
| Bidir frame | Middle third | 0.4366 | 0.1823 | |
| Bidir frame | Final third | 0.4424 | 0.09379 |