Learning via Self-Consistency for Diffusion-based Video Reasoning
Organizations: University of Toronto · University of California, Merced
Abstract
Video generation models have demonstrated emerging zero-shot capabilities for visual reasoning, perception, and other vision tasks. However, diffusion-based video generation is inherently stochastic, while many downstream vision tasks are deterministic. Motivated by the effectiveness of self-consistency in chain-of-thought reasoning for large language models, we investigate whether self-consistency can similarly improve diffusion-based video reasoning. We first introduce a training-free test-time scaling method that samples multiple video generations and aggregates their predictions through self-consistency. Specifically, we aggregate extracted paths, locations, or masks from multiple rollouts into a consensus prediction. To reduce the inference overhead of multi-rollout generation, we read out predictions early in the denoising trajectory, which preserves consensus quality while reducing denoising steps by more than half. We further propose Rejection Fine-Tuning (RFT) to distill consensus predictions into the video generation model. The resulting model internalizes the benefit of multi-sample consensus and requires only a single generation at inference time, while substantially outperforming the original model. Experiments on three tasks, including maze solving, visual search, and referring segmentation, show that both our self-consistency inference and consensus distillation dramatically improve video-based perception and reasoning, without requiring ground-truth videos or task-specific verification. For visual search, self-consistency raises task accuracy from 48.4% for a single generation to 99.0%. The distilled model retains much of the consensus benefit with a single rollout. For 4-by-4 maze solving, consensus-based training improves the single-generation strict success rate from 72.0% to 84.0% with the same inference latency.
Figures & tables
| (a) Visual search | |||
|---|---|---|---|
| Acc. | Hit | Spec. | |
| Single | 48.4 | 96.2 | 0.6 |
| Consensus | 99.0 | 98.0 | 100.0 |
| (a) Visual search | ||
|---|---|---|
| Acc. | Spec. | |
| Base | 47.8 | 2.5 |
| Consensus RFT | 69.4 | 38.8 |
| + counterfactuals | 78.4 | 56.9 |
| + contrast loss | 78.4 | 56.9 |
| Method | gIoU | Precision | Recall |
|---|---|---|---|
| Frozen-model inference | |||
| Single generation | 0.372 | 0.414 | 0.843 |
| Most-consistent sample | 0.385 | 0.423 | 0.863 |
| Pixel consensus | 0.487 | 0.570 | 0.742 |
| Student performance by target | |||
| Fixed sample (seed 0) | 0.441 | 0.484 | 0.886 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Prediction | Generations per maze | Strict validity | Goal arrival | Shortest route |
|---|---|---|---|---|
| Single generation | 1 | 14.5 | 32.4 | 11.7 |
| Modal-path vote | 128 | 50.0 | 50.0 | 50.0 |
| Method | Strict | Goal | Coverage | Shortest |
|---|---|---|---|---|
| Single generation | 59.1 | 93.1 | 100.0 | 56.7 |
| Vote | 75.6 | 84.4 | 88.9 | 75.6 |
| Vote | 62.2 | 64.4 | 66.7 | 62.2 |
| Vote | 40.0 | 42.2 | 42.2 | 40.0 |
| Vote | 24.4 | 26.7 | 26.7 | 24.4 |
| Vote | 15.6 | 17.8 | 17.8 | 15.6 |
| Method | Acc. | Hit | Precision | Spec. | Error (px) |
|---|---|---|---|---|---|
| Single generation | 48.4 | 96.2 | 49.7 | 0.6 | 7.45 |
| Vote | 51.0 | 100.0 | 50.5 | 2.0 | 0.95 |
| Vote | 61.0 | 100.0 | 56.2 | 22.0 | 0.95 |
| Vote | 79.0 | 100.0 | 70.4 | 58.0 | 0.95 |
| Vote | 88.0 | 100.0 | 80.6 | 76.0 | 0.95 |
| Vote | 91.0 | 100.0 | 84.7 | 82.0 | 0.95 |
| Method | Rollouts | Masks | gIoU | cIoU | Precision | Recall |
|---|---|---|---|---|---|---|
| Random rollout | 1 | 1 | 0.412 | 0.287 | 0.469 | 0.841 |
| Fixed seed 9 | 1 | 1 | 0.434 | 0.295 | 0.497 | 0.854 |
| Latent mode | 10 | 1 | 0.446 | 0.331 | 0.502 | 0.872 |
| Consensus (9-of-10) | 10 | 10 | 0.483 | 0.432 | 0.593 | 0.749 |
| Latent top-1 ∗ | 10 | 1 | 0.490 | 0.402 | 0.558 | 0.824 |
| Latent top-3 ∗ | 10 | 3 | 0.502 | 0.415 | 0.578 | 0.814 |
| Task / use | Rollouts | Support | Other settings |
|---|---|---|---|
| Search | 10 | 9 | Grouping radius input pixels; abstentions kept as unchanged-video targets. |
| Segmentation | 10 | 9 | Training masks kept when their foreground fraction lies in . |
| Maze inference | 10 | — | Modal path; ties broken lexicographically. |
| Maze post-training | 8 | 3 | Unique modal group required; at most three trajectories per layout. |
| Condition | Base | Consensus RFT | + counterfactuals | + contrast loss | Conf. |
|---|---|---|---|---|---|
| 2D disjunctive | 51.2 / 2.5 | 83.8 / 67.5 | 90.0 / 80.0 | 91.2 / 82.5 | 85.6 / 71.2 |
| 2D conjunctive | 40.0 / 7.5 | 56.2 / 12.5 | 67.5 / 35.0 | 66.2 / 32.5 | 57.5 / 15.0 |
| 3D disjunctive | 50.0 / 0.0 | 63.7 / 27.5 | 73.8 / 47.5 | 73.8 / 47.5 | 60.6 / 21.2 |
| 3D conjunctive | 50.0 / 0.0 | 73.8 / 47.5 | 82.5 / 65.0 | 82.5 / 65.0 | 78.8 / 57.5 |
| Model | Pass@10 (%) | Modal vote (%) |
|---|---|---|
| Base | 93.3 | 86.7 |
| Random RFT | 100.0 | 86.7 |
| Supervised RFT | 100.0 | 93.3 |
| Consensus RFT | 100.0 | 93.3 |
| 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 0.372 | — | — | — | — | — | — | — | — | — |
| 2 | 0.316 | 0.423 | — | — | — | — | — | — | — | — |
| 3 | 0.283 | 0.383 | 0.443 | — | — | — | — | — | — | — |
| 4 | 0.262 | 0.348 | 0.418 | 0.452 | — | — | — | — | — | — |
| 5 | 0.247 | 0.323 | 0.384 | 0.441 | 0.455 | — | — | — | — | — |
| 6 | 0.236 | 0.306 | 0.359 | 0.409 | 0.456 | 0.454 | — | — | — | — |
| Target | gIoU | Precision | Recall | Area |
|---|---|---|---|---|
| Fixed sample (seed 0) | 0.452 | 0.507 | 0.850 | 0.275 |
| Most-consistent sample | 0.458 | 0.499 | 0.865 | 0.276 |
| Area-controlled sample | 0.442 | 0.540 | 0.640 | 0.157 |
| Pixel consensus (9-of-10) | 0.533 | 0.627 | 0.768 | 0.157 |