DiVid: Diagnosing Dimension-Specific Diversity Collapse in Video Generation Models
Organizations: Renmin University of China · Alibaba Group · Tianjin University
Abstract
Despite remarkable progress, video generation models often produce highly similar outputs when repeatedly sampled from the same prompt, limiting their usefulness for creative exploration. Existing diversity evaluations primarily rely on global scalar metrics, which obscure where diversity collapses in the spatiotemporal space of videos. We introduce DiVid, a dimension-level diagnostic framework that decomposes video generation diversity into six interpretable dimensions: Semantic, Style, Subject, Scene, Motion, and Camera. Each dimension is measured through a reproducible computer-vision pipeline and analyzed alongside quality and instruction faithfulness to examine potential trade-offs. Systematic evaluation of representative video generation models reveals that diversity is highly dimension-specific: models with strong global diversity scores still collapse on specific factors, particularly Motion and Camera. These rankings persist after filtering unfaithful generations, indicating genuine capability differences rather than off-prompt outputs. Beyond measurement, controlled prompt interventions identify two fundamental bottlenecks: default mode convergence, where models fall back to dominant patterns under open-ended prompts; and realization gaps, where models fail to faithfully realize diverse, explicitly requested alternatives, particularly for temporal factors. The larger faithfulness losses for temporal factors highlight the difficulty of controlling motion and camera variation through text alone. DiVid thus shifts the study of diversity from measuring whether it exists to diagnosing where and why it collapses, and provides actionable directions for dimension-aware training objectives and control signals. The framework will be released to facilitate future research on diverse and controllable video generation.
Figures & tables
| Faith / Quality | Prior Metrics | Holistic | Factor Level | ||||||||||
| Model | IF | AQ | IQ | TF | MS | TCE | TIE | Semantic | Style | Subject | Scene | Motion | Camera |
| Wan2.2-5B (g=5) | 0.848 | 0.543 | 0.672 | 0.980 | 0.989 | 5.29 | 13.36 | 0.169 | 0.234 | 0.447 | 0.426 | 0.326 | 0.097 |
| Wan2.2-14B | 0.915 | 0.614 | 0.699 | 0.972 | 0.983 | 4.54 | 13.29 | 0.144 | 0.224 | 0.412 | 0.394 | 0.344 | 0.148 |
| CogVideo | 0.839 | 0.516 | 0.622 | 0.973 | 0.985 | 4.12 | 12.66 | 0.142 | 0.214 | 0.414 | 0.377 | 0.323 | 0.045 |
| Hunyuan | 0.893 | 0.541 | 0.662 | 0.965 | 0.991 | 5.05 | 13.07 | 0.142 | 0.197 | 0.392 | 0.378 | 0.433 | 0.490 |
| Wan2.7 | 0.963 | 0.563 | 0.722 | 0.972 | 0.985 | 4.94 | 12.74 | 0.158 | 0.201 | 0.397 | 0.390 | 0.413 | 0.303 |
| Model | Hinted | Omitted | Enumerated | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Subject | Scene | Motion | Camera | Scene | Motion | Camera | Subject | Scene | Motion | Camera | |
| Wan2.2-5B | +0.237 | +0.014 | +0.128 | +0.032 | +0.041 | +0.049 | +0.003 | +0.450 | +0.395 | +0.168 | +0.031 |
| Wan2.2-14B | +0.203 | +0.028 | +0.100 | +0.219 | +0.034 | -0.021 | +0.009 | +0.461 | +0.444 | +0.174 | +0.167 |
| CogVideo | +0.138 | +0.056 | +0.034 | +0.010 | +0.041 | +0.010 | +0.004 | +0.411 | +0.397 | +0.095 | +0.017 |
| Hunyuan | +0.209 | +0.075 | -0.047 | +0.081 | +0.076 | -0.020 | +0.093 | +0.492 | +0.308 | +0.067 | +0.281 |
| Wan2.7 | +0.234 | +0.109 | +0.067 | +0.398 | +0.105 | -0.004 | +0.109 | +0.468 | +0.341 | +0.146 | +0.534 |
| Dimension | Omitted | Hinted | Enumerated |
|---|---|---|---|
| Subject | N/A | an animal | {dog, cat, horse, …} |
| Scene | remove on a beach | somewhere | {beach, forest, city, …} |
| Motion | remove resting | in motion | {walks, runs, jumps, …} |
| Camera | remove static camera | camera moves | {pan, zoom, dolly, …} |
Appendix figures & tables41 assets
Supplementary material from the paper’s appendix.
Appendix
| Group | Controlled variation | Motion MPD |
|---|---|---|
| 1 | Breathing / ear twitch / deep breath | 0.254 |
| 2 | Walk left / right / toward camera | 0.469 |
| 3 | Slow / normal / fast locomotion | 0.483 |
| 4 | Sleep / stand and stretch / chase | 0.616 |
| 5 | Shake / scratch / dig | 0.456 |
| Dimension | Spearman | Kendall | Kripp. |
|---|---|---|---|
| Semantic | 0.604 | 0.539 | 0.703 |
| Style | 0.563 | 0.522 | 0.682 |
| Subject | 0.549 | 0.503 | 0.679 |
| Scene | 0.541 | 0.489 | 0.680 |
| Motion | 0.531 | 0.471 | 0.692 |
| Camera | 0.556 | 0.502 | 0.686 |
| Factor | Strict | Direction | Kripp. |
|---|---|---|---|
| Subject | 63.4% | 85.4% | 0.699 |
| Scene | 65.4% | 88.9% | 0.705 |
| Motion | 63.7% | 90.1% | 0.742 |
| Camera | 69.3% | 90.3% | 0.766 |
| Factor | VIF | Grouped-CV (95% CI) | Unexplained | Top-model change (95% CI) |
|---|---|---|---|---|
| Subject | 1.73 | 0.419 [0.363, 0.471] | 58.1% | 25.6% [19.5, 31.8] |
| Scene | 1.73 | 0.418 [0.363, 0.470] | 58.2% | 19.5% [13.8, 25.1] |
| Motion | 1.39 | 0.274 [0.222, 0.326] | 72.6% | 22.6% [16.9, 28.7] |
| Camera | 1.37 | 0.266 [0.212, 0.320] | 73.4% | 26.7% [20.5, 32.8] |
| Target factor | Prompts retrieved | Pair count |
|---|---|---|
| Subject | 122/150 (81.3%) | 286 |
| Scene | 112/150 (74.7%) | 262 |
| Motion | 138/150 (92.0%) | 484 |
| Camera | 140/150 (93.3%) | 521 |
| Model | Semantic | Style | Subject |
|---|---|---|---|
| Wan2.2-5B | |||
| Wan2.2-14B | |||
| CogVideo | |||
| Hunyuan | |||
| Wan2.7 | |||
| HappyHorse |
| Model | Scene | Motion | Camera |
|---|---|---|---|
| Wan2.2-5B | |||
| Wan2.2-14B | |||
| CogVideo | |||
| Hunyuan | |||
| Wan2.7 | |||
| HappyHorse |
| Dimension | Positive paired contrast | MPD | 95% bootstrap CI | Global BH | |
|---|---|---|---|---|---|
| Semantic | Wan2.2-5B Hunyuan | 206 | 0.027 | [0.019, 0.035] | |
| Style | Wan2.2-5B Hunyuan | 206 | 0.038 | [0.030, 0.046] | |
| Subject | Wan2.2-5B Hunyuan | 203 | 0.054 | [0.037, 0.072] | |
| Scene | Wan2.2-5B Hunyuan | 203 | 0.047 | [0.030, 0.063] | |
| Motion | Hunyuan Wan2.2-5B | 204 | 0.107 | [0.089, 0.125] | |
| Camera | Hunyuan Wan2.2-5B | 204 | 0.394 | [0.362, 0.424] |
| Source | AQ | IQ | TF | MS | Semantic | Style | Subject | Scene | Motion | Camera | Six-dim. Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Real videos | 0.574 | 0.707 | 0.980 | 0.991 | 0.239 (1) | 0.261 (1) | 0.523 (1) | 0.540 (1) | 0.407 (4) | 0.243 (4) | 0.369 (1) |
| Metric | Sem. | Sty. | Subj. | Scene | Mot. | Cam. |
|---|---|---|---|---|---|---|
| IF | -0.214 | -0.679 | -0.786 | -0.429 | 0.393 | 0.500 |
| AQ | 0.179 | -0.179 | -0.393 | -0.036 | 0.286 | 0.214 |
| IQ | -0.071 | -0.214 | -0.357 | 0.000 | 0.000 | 0.071 |
| Metric | Negative | Positive | Total |
|---|---|---|---|
| IF | 15 | 0 | 15 |
| AQ | 13 | 0 | 13 |
| IQ | 7 | 14 | 21 |
| Total | 35 | 14 | 49 |
| Model | Overall IF | Subject | Scene | Motion | Camera |
|---|---|---|---|---|---|
| Wan2.2-5B | 0.848 | 0.853 | 0.928 | 0.778 | 0.938 |
| Wan2.2-14B | 0.915 | 0.914 | 0.957 | 0.890 | 0.967 |
| CogVideo | 0.839 | 0.813 | 0.919 | 0.790 | 0.971 |
| Hunyuan | 0.893 | 0.898 | 0.937 | 0.842 | 0.977 |
| Wan2.7 | 0.963 | 0.962 | 0.992 | 0.935 | 0.986 |
| HappyHorse | 0.986 | 0.984 | 0.999 | 0.968 | 1.000 |
| Faith / Quality | Prior Metrics | Holistic | Factor Level | ||||||||||
| Model | IF | AQ | IQ | TF | MS | TCE | TIE | Semantic | Style | Subject | Scene | Motion | Camera |
| Wan2.2-5B | 0.845 | 0.553 | 0.682 | 0.982 | 0.990 | 5.555 | 13.422 | 0.184 | 0.241 | 0.473 | 0.450 | 0.317 | 0.083 |
| Wan2.2-14B | 0.916 | 0.620 | 0.704 | 0.972 | 0.984 | 4.590 | 13.262 | 0.149 | 0.228 | 0.425 | 0.406 | 0.338 | 0.146 |
| CogVideo | 0.857 | 0.531 | 0.636 | 0.974 | 0.986 | 4.006 | 12.385 | 0.148 | 0.217 | 0.421 | 0.382 | 0.315 | 0.051 |
| Hunyuan | 0.899 | 0.549 | 0.670 | 0.965 | 0.991 | 5.318 | 13.263 | 0.154 | 0.206 | 0.416 | 0.399 | 0.450 | 0.504 |
| Wan2.7 | 0.965 | 0.572 | 0.731 | 0.972 | 0.985 | 5.177 | 12.765 | 0.170 | 0.206 | 0.414 | 0.408 | 0.422 | 0.308 |
| Faith / Quality | Prior Metrics | Holistic | Factor Level | ||||||||||
| Model | IF | AQ | IQ | TF | MS | TCE | TIE | Semantic | Style | Subject | Scene | Motion | Camera |
| Wan2.2-5B | 0.858 | 0.518 | 0.645 | 0.974 | 0.986 | 4.564 | 13.198 | 0.131 | 0.216 | 0.376 | 0.363 | 0.349 | 0.134 |
| Wan2.2-14B | 0.912 | 0.598 | 0.683 | 0.970 | 0.981 | 4.388 | 13.368 | 0.131 | 0.211 | 0.376 | 0.361 | 0.358 | 0.155 |
| CogVideo | 0.789 | 0.475 | 0.585 | 0.970 | 0.983 | 4.418 | 13.407 | 0.127 | 0.207 | 0.395 | 0.363 | 0.345 | 0.030 |
| Hunyuan | 0.876 | 0.517 | 0.641 | 0.965 | 0.991 | 4.351 | 12.546 | 0.109 | 0.173 | 0.329 | 0.323 | 0.389 | 0.453 |
| Wan2.7 | 0.958 | 0.539 | 0.700 | 0.972 | 0.985 | 4.310 | 12.664 | 0.129 | 0.187 | 0.353 | 0.343 | 0.390 | 0.289 |
| Dimension | Most complementary pair | |
|---|---|---|
| Semantic | CogVideo + HappyHorse | 0.140 |
| Style | CogVideo + HappyHorse | 0.117 |
| Subject | CogVideo + HappyHorse | 0.230 |
| Scene | CogVideo + HappyHorse | 0.257 |
| Motion | Wan2.2-5B + Seedance | 0.139 |
| Camera | CogVideo + Seedance | 0.152 |
| Setting | Construction | Diagnostic role |
|---|---|---|
| Specified | All four dimensions are explicitly specified. | Paired reference condition. |
| Omitted | Remove the target-dimension description. | Tests autonomous exploration. |
| Hinted | Replace the specific description with a generic cue. | Tests dimension activation. |
| Enumerated | Enumerate target candidates; generate one video per prompt. | Tests candidate realization. |
| Setting | Target | No. | Prompt |
| Specified | All | – | A large fluffy adult golden retriever is resting on a quiet sandy beach at golden hour, static camera |
| Hinted | Subject | – | An animal is resting on a quiet sandy beach at golden hour, static camera |
| Hinted | Scene | – | A large fluffy adult golden retriever is resting somewhere, static camera |
| Hinted | Motion | – | A large fluffy adult golden retriever is in motion on a quiet sandy beach at golden hour, static camera |
| Hinted | Camera | – | A large fluffy adult golden retriever is resting on a quiet sandy beach at golden hour, camera moves |
| Omitted | Scene | – | A large fluffy adult golden retriever is resting, static camera |
| Intervention | Dimension | Mean [95% CI] | Positive blocks | |||
|---|---|---|---|---|---|---|
| Hinted | Subject | +0.197 [+0.144,+0.251] | 19/20 | 1.49 | ||
| Hinted | Scene | +0.056 [+0.032,+0.087] | 17/20 | 0.81 | ||
| Hinted | Motion | +0.069 [+0.044,+0.094] | 18/20 | 1.13 | ||
| Hinted | Camera | +0.210 [+0.184,+0.239] | 20/20 | 3.05 | ||
| Omitted | Scene | +0.051 [+0.025,+0.082] | 18/20 | 0.73 | ||
| Omitted | Motion | +0.011 [-0.007,+0.029] | 12/20 | 0.26 | 0.139 | 0.139 |
| Contrast | Dimension | Mean difference [95% CI] | Positive blocks | Positive models | |
|---|---|---|---|---|---|
| Hinted Omitted | Scene | +0.005 [-0.007,+0.018] | 10/20 | 4/7 | 0.298 |
| Hinted Omitted | Motion | +0.058 [+0.032,+0.087] | 17/20 | 6/7 | |
| Hinted Omitted | Camera | +0.118 [+0.092,+0.147] | 20/20 | 6/7 | |
| Enumerated Hinted | Subject | +0.275 [+0.250,+0.301] | 20/20 | 7/7 | |
| Enumerated Hinted | Scene | +0.313 [+0.262,+0.364] | 20/20 | 7/7 | |
| Enumerated Hinted | Motion | +0.077 [+0.057,+0.098] | 20/20 | 7/7 |
| Dimension | Hinted mean | Enumerated mean | Enumerated Hinted [95% CI] | Negative blocks | Negative models | |
|---|---|---|---|---|---|---|
| Subject | 4.593 | 4.497 | -0.096 [-0.258,+0.048] | 10/20 | 5/7 | 0.226 |
| Scene | 4.999 | 4.791 | -0.204 [-0.254,-0.156] | 19/20 | 7/7 | |
| Motion | 4.780 | 4.090 | -0.690 [-0.887,-0.486] | 18/20 | 7/7 | |
| Camera | 4.469 | 4.075 | -0.393 [-0.541,-0.247] | 18/20 | 6/7 |