We study whether scaling video generation enables models to reason about hidden information from the past frames, and at what computational cost. Our controlled benchmark requires predicting nine prescribed moves of an initially solved 2x2x2 Rubik's Cube from a fixed view of three faces. Correct predictions require inferring how actions change hidden states, and the simulator provides exact ground truth for evaluation. Models learn plausible cube geometry early, while correct sticker configurations require substantially more training. Although validation MSE follows approximate power-law scaling, lower MSE loss does not reliably indicate downstream reasoning capabilities. Smaller autoregressive models achieve higher state accuracy with limited compute, while larger models reach higher accuracy after more training. At roughly 0.1 PF-days, the 70M-parameter model correctly predicts the visible sticker configuration in 44.6% of post-action frames, compared with 0.3% for the 1B model, which reaches 83.7% at 3.14 PF-days. Symbolic state supervision raises the 20M model's frame accuracy from 31.1% to 67.3% at the same training-data budget, suggesting that learning representations of state changes can complement scaling.
Figures & tables
Figure 1: Rubik video generation and scaling with training compute. The video strip shows 80 predicted frames, or 3.3 s at 24 fps, for 9 prescribed actions as the language prompt from a solved cube . The video generator is trained from scratch on synthetic videos. (a) Validation flow MSE and (b) average frame accuracy for AR models of different sizes. (c) AR full-trajectory accuracy, which requires all nine post-action states to be correct. Separate power-law fits to the middle and final thirds of its observed compute frontier illustrate how compute demand grows toward higher reliability. Diamonds mark their window-dependent 98% intersections.
Figure 2: Predictions from the 270M-parameter autoregressive model after four prescribed moves from a solved cube . The model reproduces the cube’s geometry after 3.4×10−2 PF-days of training, but first predicts all 12 visible stickers correctly in this example at 5.7×10−1 PF-days, requiring about 17 times as much compute. The top-left image is the ground truth.
Figure 3: Rubik face-turn notation. Orange cubies rotate around the blue axes. Arrows show 90∘ clockwise turns viewed from outside the named face. Dashed arrows mark hidden faces. A prime reverses the turn, and a suffix 2 denotes 180∘ .
Metric
Definition
MSE loss
Mean squared error between the predicted and target flow velocities in video latent space on validation set.
Action following acc.
Fraction of actions in the prompt that are correctly applied to the cube.
Sticker acc.
Fraction of visible stickers ( 4 stickers ×3 visible faces) with the correct color after each action.
Frame acc.
Fraction of the nine post-action frames in which all 12 visible stickers have the correct color.
Full-trajectory acc.
Fraction of episodes in which all 12 visible stickers are correct after every one of the nine actions.
Table 1: Evaluation metrics.
Configuration
dmodel
Layers
Heads
dff
Parameters (M)
Generation
Action
Sticker
Frame
Deep-narrow
512
22
8
1408
95.03
AR- k=1
17.22
25.37
3.67
AR- k=4
6.33
22.20
0.78
Bidir
7.33
21.59
0.33
Balanced
640
14
10
1728
94.20
AR- k=1
72.44
52.00
24.33
AR- k=4
9.67
22.46
1.67
Bidir
6.78
21.55
0.56
Table 2: Stage A architectures and results after 1M training videos. Action, sticker, and frame accuracies (%) use the same 100 held-out videos. AR- k generates k latent frames per chunk. Best accuracies are bold.
Figure 4: Video latent space and vector space. The action prompt is omitted. (a) A frozen Wan VAE encodes video frames into continuous latents for the video DiT. (b) Visible sticker colors are represented directly as 12×6 one-hot arrays for the state Transformer. Both model families support autoregressive and bidirectional generation.
Generation
PF-days (approx.)
Frame accuracy after 20 actions (%)
Full-trajectory accuracy (%)
AR
8×10−3
94.01
92.81
Bidir
4×10−4
99.86
99.85
Table 3: Vector-space prediction on 10,000 held-out 20-action episodes. Compute estimates exclude the frozen language encoder. Final frame accuracy requires all 12 visible stickers to be correct after 20 actions. Full-trajectory accuracy requires them to be correct after every action. AR uses its own predictions during rollout.
Figure 5: State prediction with three or six visible faces. (a) Ground truth and generated six-face videos after each action in the same 9-action sequence. Red boxes mark each model’s first incorrect boundary frame. (b) Frame accuracy after each action, with 95% normal confidence intervals over 100 paired episodes. Training exposure is matched between 3-face and 6-face. Frame accuracy requires all 12 visible stickers to be correct for three-face observation and all 24 for six-face observation.
Figure 6: Frame accuracy after each action as training data increases, for AR (top) and Bidir (bottom). Colors indicate training videos on a shared scale. All selected checkpoints use the same 100 free-running episodes. Smooth lines connect the measured action accuracies.
Figure 7: Illustrative compute extrapolations for (a) AR average frame accuracy, (b) AR full-trajectory accuracy, and (c) Bidir average frame accuracy. Gray points show evaluated checkpoints; colored points identify the middle and final thirds of each compute-ordered frontier. Solid lines show fits within each segment, and dashed lines show extrapolations. Diamonds mark 98% accuracy, labeled in PF-days. Shading marks the observed compute range. Bidir full-trajectory accuracy is omitted because it is zero at every evaluated checkpoint.
Figure 8: Validation MSE and reasoning accuracy across model sizes, for AR (top) and Bidir (bottom). Points pair validation MSE with free-running accuracy from the same EMA checkpoint. Colors identify model sizes.
Figure 9: Explicit state guidance for 20M AR- k=1 models. (a) A shared Transformer predicts sticker-state distributions from the available video history and action prompt, then uses them to condition a separate video pass. State cross-entropy and video flow matching jointly train the model. (b) Frame accuracy after each action on the same 100 paired episodes after training on 3M videos. All curves use generated history. Shading shows 95% confidence intervals. All models use EMA weights.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Video
Sticker
Frame
Full trajectory
Original RGB
100.00
100.00
100.00
Wan VAE reconstruction
100.00
100.00
100.00
Appendix
Table 4: State accuracy (%) on the same 100 paired evaluation videos before and after reconstruction with the frozen Wan VAE. All metrics use the nine post-action boundaries and exclude the initial frame.
Source
Action following (all)
Action following (A9)
VAE reconstruction
92.6
92.1
AR
80.1
80.6
Bidirectional
69.6
67.7
Appendix
Table 5: Rubik action following from grayscale motion. The predicted face and turn type must both match the prompt. Values are percentages over 9,000 action segments, except the action-nine column, with 1,000 segments.
Figure 10: Vector-state learning curves for Rubik’s Cube.
Component
Value or rule
Scaling treatment
Transformer block
Self-attention, cross-attention, SwiGLU
Fixed
Model width dmodel
Tables 2 and 7
Scaled
Depth L
Selected wide–shallow ladder
Scaled
Position encoding
3D RoPE
Fixed
Head width dhead
64
Fixed
Attention heads H
dmodel/64
Derived
Appendix
Table 6: Video DiT architecture used in the two-stage scaling study.
Model name
Parameters (M)
dmodel
Layers
Heads
dff
20M
19.70
384
8
6
1024
30M
33.45
448
10
7
1216
70M
67.81
640
10
10
1728
120M
115.8
768
12
12
2048
270M
271.8
1088
14
17
2944
550M
548.1
1408
17
22
3776
Appendix
Table 7: Stage B Video DiT architecture specifications. AR- k=1 and bidirectional models share the same architecture dimensions and parameter counts.
Generation
Peak learning rate
10−4
2×10−4
4×10−4
AR- k=1
0.1678
0.1217
0.0768
AR- k=2
0.1790
0.1337
0.0810
AR- k=4
0.1715
0.1287
0.0783
Bidir
0.1789
0.1274
0.0781
Appendix
Table 8: Validation flow MSE in the learning-rate (left) and effective-batch-size (right) searches. Lower is better; best values within each comparison are bold.
Shape
AR- k=1
AR- k=4
Bidir
Deep–narrow
5.015
5.281
6.343
Balanced
4.542
4.753
5.527
Wide–shallow
4.367
4.548
5.150
Appendix
Table 9: Estimated Stage-A training compute , in 1018 FLOPs, after 1M videos. One multiply–accumulate counts as two FLOPs; backward compute is approximated as twice forward compute. These are operation-count estimates, not hardware timings.
Shape
Family
MSE ( 10−3 )
Action
Sticker
Frame
deep–narrow
AR-K1
2.233
17.22
25.37
3.67
deep–narrow
AR-K4
2.562
6.33
22.20
0.78
deep–narrow
BIDIR
3.212
7.33
21.59
0.33
balanced
AR-K1
1.883
72.44
52.00
24.33
balanced
AR-K4
2.533
9.67
22.46
1.67
balanced
BIDIR
3.426
6.78
21.55
0.56
Appendix
Table 10: Stage A after 1M training videos. Accuracies (%) average actions 1–9 on the same 100 episodes, excluding the initial boundary. Flow MSE uses 1,024 validation episodes; its conditioning and fixed probes differ across prediction factorizations.
Figure 11: Stage-A calibration across nine setups. Purple, teal, and orange group deep–narrow, balanced, and wide–shallow; solid, dashed, and dotted lines distinguish AR- k=1 , AR- k=4 , and Bidir. (a–c) The 1M-video EMA checkpoints on 100 paired development episodes: action-following accuracy from the independent motion probe, visible-sticker accuracy, and exact-frame accuracy (all 12 stickers correct). The initial boundary is excluded. (d) EMA flow MSE on 256 fixed episodes at seven selected exposures: 32,768, 65,536, 131,072, 262,144, 524,288, 786,432, and 1M videos. Diamond endpoint markers use 1,024 episodes. Initialization is omitted. AR loss uses ground-truth history; conditioning and noise/time probes differ across factorizations, so their losses are not directly comparable.
Figure 12: Stage-A rollouts across all nine setups after 1M training videos. VAE-reconstructed ground truth, followed by deep–narrow, balanced, and wide–shallow; each shape includes AR- k=1 , AR- k=4 , and Bidir. All use the first development episode (seed 81024000), the same prompt, and paired sampling-noise seeds, with 16 midpoint steps per AR chunk or entire Bidir video. The nine settled post-action boundaries are shown. Red boxes mark each model’s first incorrect settled action boundary. The episode was not selected by performance. All panels use the saved EMA model parameters.
Model
Mean frame: Free
GT prefix
A9 frame: Free
GT prefix
XS
33.89
38.67
0.00
5.00
S
42.22
47.89
0.00
9.00
M
52.56
60.00
0.00
22.00
B
48.11
55.00
1.00
9.00
L
63.56
71.67
3.00
38.00
XL
64.56
73.33
5.00
45.00
Appendix
Table 11: Matched free-running and action-aligned GT-history comparisons at 5M videos. All values are percentages on the same 100 episodes.
Operation
Parameters
Forward FLOPs per target token
Language projection
dfdmodel
2ntextdfdmodel/ntgt
Latent input/output projections
3dzp2dmodel
2dzp2dmodel(2+ntgtnanchor+nhist)
Target visual Q/K/V and output
4Ldmodel2
8Ldmodel2
Visual-context K/V (shared weights)
—
4Lntgtnanchor+nhistdmodel2
Visual attention products
—
4L(ntgt+nanchor+nhist)dmodel
Language cross-attention projections
4Ldmodel2
4Ldmodel2(1+ntgtntext)
Appendix
Table 12: Generator parameters and forward compute. FLOPs are per target visual patch token; biases, normalization, and nonlinearities are omitted.
Metric
Segment
A
α
C98
AR frame
Middle third
0.336
0.2178
4.227×105
AR frame
Final third
0.5607
1.055
23.55
AR trajectory
Middle third
0.9501
0.1185
1.425×1014
AR trajectory
Final third
1.59
0.7931
248.9
Bidir frame
Middle third
0.4366
0.1823
2.221×107
Bidir frame
Final third
0.4424
0.09379
2.184×1014
Appendix
Table 13: Power-law fits to the middle and final thirds of each observed compute frontier. C98 is the 98% intersection of each fitted curve, in PF-days, illustrating its sensitivity to the fitting window.