Video generation models have demonstrated emerging zero-shot capabilities for visual reasoning, perception, and other vision tasks. However, diffusion-based video generation is inherently stochastic, while many downstream vision tasks are deterministic. Motivated by the effectiveness of self-consistency in chain-of-thought reasoning for large language models, we investigate whether self-consistency can similarly improve diffusion-based video reasoning. We first introduce a training-free test-time scaling method that samples multiple video generations and aggregates their predictions through self-consistency. Specifically, we aggregate extracted paths, locations, or masks from multiple rollouts into a consensus prediction. To reduce the inference overhead of multi-rollout generation, we read out predictions early in the denoising trajectory, which preserves consensus quality while reducing denoising steps by more than half. We further propose Rejection Fine-Tuning (RFT) to distill consensus predictions into the video generation model. The resulting model internalizes the benefit of multi-sample consensus and requires only a single generation at inference time, while substantially outperforming the original model. Experiments on three tasks, including maze solving, visual search, and referring segmentation, show that both our self-consistency inference and consensus distillation dramatically improve video-based perception and reasoning, without requiring ground-truth videos or task-specific verification. For visual search, self-consistency raises task accuracy from 48.4% for a single generation to 99.0%. The distilled model retains much of the consensus benefit with a single rollout. For 4-by-4 maze solving, consensus-based training improves the single-generation strict success rate from 72.0% to 84.0% with the same inference latency.
Figures & tables
Figure 1: Self-consistency improves video reasoning for inference and post-training. (Left) Consensus solution is better than individual prediction and supports early readout and distillation. (Right) Inference with self-consistency improves MiniMax H3 predictions through aggregation; learning with self-consistency turns consensus into faster and sometimes stronger single-generation predictions.
Figure 2: Consensus teacher construction and student training. (a) A frozen video model generates multiple samples, whose extracted paths, points, or masks are aggregated in task space. Optional early readout decodes a predicted clean latent from a prefix of the native denoising schedule. (b) Consensus masks and points become rendered target videos, including unchanged inputs for abstention, and are VAE-encoded into fixed targets. Mazes instead reuse clean latents from consensus-selected trajectories. Only the student’s LoRA parameters are optimized. Thumbnails use real inputs and generated frames; reveal strips illustrate rendered consensus targets.
Figure 3: Maze consensus and trajectory selection on an illustrative input. Actual video frames are overlaid with extracted cell paths. Eight sampled paths form five agreement groups. The unique modal group contains seeds 0, 1, and 7 (three of eight), whose generated trajectories are reused as training targets for RFT.
(a) Visual search
Acc.
Hit
Spec.
Single
48.4
96.2
0.6
Consensus
99.0
98.0
100.0
Table 1: Single generation versus ten-seed consensus using MiniMax-H3 with the same inputs and seeds. By default, we use 49 steps during inference unless marked w/ early (with early readout). Search and maze scores are percentages; time is in seconds. All metrics except time are higher-is-better. Calls count denoiser calls. Best results are in bold.
(a) Visual search
Acc.
Spec.
Base
47.8
2.5
Consensus RFT
69.4
38.8
+ counterfactuals
78.4
56.9
+ contrast loss
78.4
56.9
Table 2: Single-generation performance (w/ 20 denoising steps) after our consensus-based RFT. Search and maze scores are percentages; hole and jump rates are lower-is-better. Area is the predicted foreground fraction. Random RFT is to fine-tune with random rollouts. Supervised RFT uses a verifier to verify rollouts and select correct ones as supervision.
Figure 4: Referring-segmentation outputs for the prompt “jeep just to right of clown”. Base and trained single-generation panels show final frames with the same seed. The consensus panel visualizes a three-of-four vote over the base model’s extracted masks; all outputs use early readout. The trained panel uses the model after consensus RFT from Table 2 .
Method
gIoU ↑
Precision ↑
Recall ↑
Frozen-model inference
Single generation
0.372
0.414
0.843
Most-consistent sample
0.385
0.423
0.863
Pixel consensus
0.487
0.570
0.742
Student performance by target
Fixed sample (seed 0)
0.441
0.484
0.886
Table 3: Segmentation ablations on separate inference and student-evaluation cohorts. Top: aggregation of ten frozen-model predictions per image. Bottom: students trained on the same 211 inputs with the same three-epoch recipe. Student scores average four single-generation draws per image; all predictions use 20-call sampling.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Prediction
Generations per maze
Strict validity
Goal arrival
Shortest route
Single generation
1
14.5
32.4
11.7
Modal-path vote
128
50.0
50.0
50.0
Appendix
Table 4: Wan2.2 results on two Frozen Lake mazes (%). Single-generation scores average all 128 samples per maze. Voting returns the most frequent path for each maze. Scores average over the two mazes; a voting success rate of 50.0% means one of the two mazes is solved.
Method
Strict
Goal
Coverage
Shortest
Single generation
59.1
93.1
100.0
56.7
Vote ≥2/10
75.6
84.4
88.9
75.6
Vote ≥3/10
62.2
64.4
66.7
62.2
Vote ≥4/10
40.0
42.2
42.2
40.0
Vote ≥5/10
24.4
26.7
26.7
24.4
Vote ≥6/10
15.6
17.8
17.8
15.6
Appendix
Table 5: Maze inference on 45 layouts with 10 seeds each (%). Unanswered inputs count as failures.
Method
Acc.
Hit
Precision
Spec.
Error (px)
Single generation
48.4
96.2
49.7
0.6
7.45
Vote ≥2/10
51.0
100.0
50.5
2.0
0.95
Vote ≥3/10
61.0
100.0
56.2
22.0
0.95
Vote ≥4/10
79.0
100.0
70.4
58.0
0.95
Vote ≥5/10
88.0
100.0
80.6
76.0
0.95
Vote ≥6/10
91.0
100.0
84.7
82.0
0.95
Appendix
Table 6: Visual search on 100 arrays with 10 seeds each. Rates are percentages; error is the localization error in input pixels. The 9-of-10 rule is the one used throughout the paper; the other thresholds are shown for reference.
Figure 5: Segmentation gIoU against (a) the number of seeds, (b) denoiser calls, and (c) inference time on the 100 segmentation inputs. Each value averages every subset of size G , with support m=⌈0.9G⌉ . Early readout (20 steps) matches full generation (49 steps) at every number of seeds and reaches the same quality with fewer calls and less time.
Method
Rollouts
Masks
gIoU
cIoU
Precision
Recall
Random rollout
1
1
0.412
0.287
0.469
0.841
Fixed seed 9
1
1
0.434
0.295
0.497
0.854
Latent mode
10
1
0.446
0.331
0.502
0.872
Consensus (9-of-10)
10
10
0.483
0.432
0.593
0.749
Latent top-1 ∗
10
1
0.490
0.402
0.558
0.824
Latent top-3 ∗
10
3
0.502
0.415
0.578
0.814
Appendix
Table 7: Choosing rollouts from early latents on a second set of 100 segmentation inputs. All rollouts use early readout (20 steps). “Masks” counts the decoded masks used for the final answer. ∗ Uses a supervised latent selector.
Task / use
Rollouts G
Support m
Other settings
Search
10
9
Grouping radius r=25 input pixels; abstentions kept as unchanged-video targets.
Segmentation
10
9
Training masks kept when their foreground fraction lies in [0.002,0.6] .
Maze inference
10
—
Modal path; ties broken lexicographically.
Maze post-training
8
3
Unique modal group required; at most three trajectories per layout.
Appendix
Table 8: Consensus and target settings.
Figure 6: A counterfactual training pair. The object at the teacher’s consensus point is recolored, and the teacher abstains on the edited array, which gives an unchanged-video target. The circle marks the edited object and is not part of the input.
Condition
Base
Consensus RFT
+ counterfactuals
+ contrast loss
Conf.
2D disjunctive
51.2 / 2.5
83.8 / 67.5
90.0 / 80.0
91.2 / 82.5
85.6 / 71.2
2D conjunctive
40.0 / 7.5
56.2 / 12.5
67.5 / 35.0
66.2 / 32.5
57.5 / 15.0
3D disjunctive
50.0 / 0.0
63.7 / 27.5
73.8 / 47.5
73.8 / 47.5
60.6 / 21.2
3D conjunctive
50.0 / 0.0
73.8 / 47.5
82.5 / 65.0
82.5 / 65.0
78.8 / 57.5
Appendix
Table 9: Search accuracy / specificity (%) by condition. Development: 20 arrays per condition, four seeds each. Conf.: Consensus RFT on the 160 confirmation arrays (40 per condition).
Model
Pass@10 (%)
Modal vote (%)
Base
93.3
86.7
Random RFT
100.0
86.7
Supervised RFT
100.0
93.3
Consensus RFT
100.0
93.3
Appendix
Table 10: More 4×4 maze results on the 60 held-out layouts (20 per hole density, 10 seeds each). Pass@10 counts a layout as solved when any of its ten generations is valid. The modal vote picks the most frequent path without checking validity.
G\m
1
2
3
4
5
6
7
8
9
10
1
0.372
—
—
—
—
—
—
—
—
—
2
0.316
0.423
—
—
—
—
—
—
—
—
3
0.283
0.383
0.443
—
—
—
—
—
—
—
4
0.262
0.348
0.418
0.452
—
—
—
—
—
—
5
0.247
0.323
0.384
0.441
0.455
—
—
—
—
—
6
0.236
0.306
0.359
0.409
0.456
0.454
—
—
—
—
Appendix
Table 11: Early-readout segmentation gIoU for G samples and support m . Bold entries mark the rule m=⌈0.9G⌉ .
Target
gIoU ↑
Precision ↑
Recall ↑
Area
Fixed sample (seed 0)
0.452
0.507
0.850
0.275
Most-consistent sample
0.458
0.499
0.865
0.276
Area-controlled sample
0.442
0.540
0.640
0.157
Pixel consensus (9-of-10)
0.533
0.627
0.768
0.157
Appendix
Table 12: Quality of the segmentation training targets, measured against ground truth. Area is the foreground fraction.
Figure 7: Pixel votes over 10 seeds for one referring expression. Color shows how many of the ten extracted masks cover each pixel.
Figure 8: Consensus v.s. single, and full generation (49 steps) v.s. early readout (20 steps). “Consensus” is the 9-of-10 vote over 10 seeds. Numbers are IoU w.r.t. ground truth.
Figure 9: Referring segmentation before and after consensus RFT on testing images (seed 0, early readout). Numbers are IoU with the ground truth.
Figure 10: Maze solving before (top) and after (bottom) consensus RFT on held-out 4×4 layouts. Each column uses the same layout and seed. Lines show the tracked cell path, drawn over the last generated frame.
Figure 11: Visual search before and after consensus RFT on development arrays, with the same seed in each row. Rings mark detected blue markers; a panel without a ring has no marker.
Video diffusion models have made rapid progress in perceptual realism and temporal coherence, but they remain primarily optimized for plausible generation rather than verifiable reasoning. This limitation is especially pronounced in tasks where generated videos must satisfy explicit spatial, temporal, or logical constraints. Inspired by the role of reinforcement learning with verifiable rewards (RLVR) in reasoning-oriented language models, we introduce VideoRLVR, a practical recipe for optimizing video diffusion models with rule-based feedback. VideoRLVR formulates video reasoning as the generation of verifiable visual trajectories and consists of an SDE-GRPO optimization backbone, dense decomposed rewards, and an Early-Step Focus strategy for efficient training. The Early-Step Focus strategy restricts policy optimization to the early denoising phase, reducing training latency by about 40% while preserving performance. We evaluate VideoRLVR on Maze, FlowFree, and Sokoban, three procedurally generated domains with objective success criteria. Across these tasks, VideoRLVR consistently improves over supervised fine-tuning baselines, with dense decomposed rewards proving especially important in low-success-rate settings. Our RL-optimized model also outperforms the evaluated proprietary and open-source video generation models on these verifiable reasoning benchmarks and out-of-domain benchmarks. These results suggest that verifiable RL can move video models beyond perceptual imitation toward more reliable rule-consistent visual reasoning.
Tinghui Zhu, Sheng Zhang, James Y. Huang +5
University of California, Davis · Microsoft Research · University of Southern California +1
While test-time scaling has revolutionized reasoning in large language models, generative video reasoning remains bottlenecked by a single-shot paradigm. We demonstrate that searching over denoising steps cannot rescue logically flawed rollouts because spatial trajectories commit early in the diffusion process. Root-level Best-of-N (BoN) sampling is similarly inefficient: reasoning errors cluster early in the temporal axis, and resampling blindly discards verified upstream progress. To unlock effective test-time scaling for video models, we introduce Temporal Backtracking Search (TBS), which shifts the search space to the temporal axis. TBS transforms video generation into an iterative generate-verify-restart loop via three core mechanisms: (1) variable-K conditioning to resume generation from arbitrary clean prefixes; (2) temporal process verification to localize failures and extract valid restart anchors; and (3) prefix-based search to reallocate compute toward extending correct trajectories rather than root resampling. Across algorithmic, navigation, and robotics domains, TBS Pareto-dominates matched-budget BoN. In a strict out-of-distribution setting where one-shot generation collapses (0.7% for BoN), TBS achieves 22.7%, with every solved episode stemming from a restarted branch. Ultimately, TBS reveals that the local reasoning competence of video models far exceeds what single-shot rollouts indicate, providing a scalable test-time framework to unlock it.
Sejoon Jun, Zheng Ding, Huangyuan Su +2
Northeastern University · Independent Researcher · Harvard & Kempner +1
Causal video generators must predict from the past, but they need not learn only from it. In streaming autoregressive video diffusion, each emitted segment becomes a commitment that future segments must preserve. Standard training, however, only asks each causal state to explain the present. This creates what we call a representation-level planning gap: states that fit the current segment may discard identity, layout, and motion information needed for a consistent future. We introduce Video-Mirai, a training-only method that closes this gap without changing causal inference: the generator rolls out causally, a frozen foresight encoder reads the completed rollout non-causally, and a lightweight predictor distills the resulting stopped-gradient targets into causal states. Future frames supervise representations, never generator inputs. At inference, the encoder and predictor are discarded, leaving the original architecture, per-step FLOPs, and KV-cache behavior unchanged. Video-Mirai improves a strong Causal-Forcing baseline on 5-second VBench from 83.8 to 84.6 in terms of Total Score. On 30-second rollouts beyond the training horizon, subject consistency improves from 84.9 to 88.5 and background consistency from 90.2 to 91.9. Ablations identify future-conditioned targets as the key ingredient, and probes show that future frames become more decodable from current features. Causality should constrain inference, not representation supervision. Our study highlights that visual autoregressive models need foresight. Project page: https://y0uroy.github.io/Video-Mirai.
Yonghao Yu, Lang Huang, Runyi Li +2
The University of Tokyo · National Institute of Informatics · Peking University