Evaluating semantic similarity between videos is a fundamental challenge in computer vision, essential for tasks ranging from out-of-distribution (OOD) detection to video retrieval. However, defining and labeling video similarity is notoriously difficult and expensive due to the complex spatio-temporal nature. In this paper, we propose a novel self-supervised approach that leverages generative uncertainty from text-to-video (T2V) diffusion models to learn semantic similarity without human annotations. Our method is based on the observation that T2V models produce consistent outputs for familiar concepts but exhibit high variance and uncertainty when prompted with specialized concepts. We utilize this behavior to identify stable semantic features within existing pretrained representations, such as VideoMAE and V-JEPA. Specifically, we learn a mask over these embeddings using purely generated data, encouraging the model to retain features that remain consistent across generations of general concepts while discarding those associated with generative noise or uncertainty. Experimental results across three key tasks demonstrate that our learned feature subspaces consistently outperform original pretrained features and baseline feature selection methods.
Figures & tables
Figure 1: Proposed framework overview. (1) A text-to-video model generates videos from general and specialized prompts with different random seeds; videos from the same prompt are paired. (2) A frozen video encoder extracts global representations, and a learnable mask is trained to yield high/low cosine similarity for similar/dissimilar pairs. (3) For downstream tasks, we mask unseen video representations before cosine similarity. In doing so, our strategy improves alignment with human judgment, out-of-distribution detection, and video retrieval.
Figure 2: Example videos used in our generative consistency objective. Videos generated by Wan 2.1 with general and specialized text prompts. Videos from general prompts serve as proxies for semantic consistency, while those from specialized prompts represent generative divergence.
Figure 3: Per-prompt intra-set feature-space diversity for videos generated by Wan2.1, measured using V-JEPA 2.1 Large (left) and VideoMAEv2 Base (right). Each point represents one prompt and reports the average pairwise normalized cosine dissimilarity among three videos generated with different random seeds.
Main Action
Main Subjects
Main Objects
Location
Actions Order
Average
Encoder
Method
Sp. (%)
ρ
τ
ρ
τ
ρ
τ
ρ
τ
ρ
τ
ρ
τ
Baseline (no mask)
–
24.0
16.30
33.4
22.84
24.7
16.83
49.3
35.04
24.3
16.65
31.14
21.53
Binary mask
67.8
26.9
18.25
39.4
27.14
30.3
20.57
50.5
35.74
30.1
20.75
35.44
24.49
V-JEPA 2.1-L
Soft mask
–
27.7
18.78
39.5
27.24
30.8
20.89
51.3
36.41
30.9
21.23
36.04
24.91
Baseline (no mask)
–
47.0
33.02
44.6
31.46
42.4
29.96
45.8
32.26
44.7
31.37
44.90
31.61
Binary mask
58.0
47.9
33.62
46.3
32.63
44.1
30.98
46.9
33.09
45.3
31.67
46.10
32.40
Table 1: Alignment with human judgement on ConViS-Bench. Spearman’s ρ and Kendall’s τ correlations with human judgement ( ×100 ). Sp. denotes the sparsity of the learned mask; soft masks are dense by construction. Our masked features better match the human notion of global video similarity. Best results are bolded , second-best underlined .
V-JEPA 2.1-L
VideoMAEv2-B
2XDM
DDPM-OOD
2XDM
DDPM-OOD
OOD Dataset
Method
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
UVEB
Unmasked
78.42
73.77
57.20
76.37
70.31
63.60
71.08
68.50
77.60
75.50
73.33
83.20
Binary mask
86.77
85.72
51.20
89.52
88.04
42.80
72.31
69.27
73.60
77.47
75.10
84.80
Soft mask
85.79
84.48
52.00
89.11
87.35
42.80
72.21
68.90
73.20
77.37
74.81
83.20
MedVidBench
Unmasked
74.66
74.18
76.00
68.11
68.47
86.00
78.59
80.73
82.40
77.01
79.46
79.60
Table 2: Video generation OOD detection across four diverse datasets. Videos are generated using Cosmos-Predict2, with BDD100k as the in-distribution (ID) baseline. Applying binary and soft feature masks consistently improves OOD detection performance. Best results are bolded and second-best results are underlined .
Dataset
Encoder
Retrieval mAP (%)
Unmasked
Binary mask
Soft mask
UCF101
V-JEPA 2.1 Large
40.82
41.34
42.37
VideoMAEv2-Base
95.52
95.61
95.65
HMDB51
V-JEPA 2.1 Large
17.84
18.40
18.99
VideoMAEv2-Base
46.94
46.39
46.66
Table 3: Class-level video retrieval mAP (%) on UCF101 and HMDB51. Best results are bolded and second-best results are underlined .
V-JEPA 2.1-L
VideoMAEv2-B
2XDM
DDPM-OOD
2XDM
DDPM-OOD
Method
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
Baseline (unmasked)
71.75
66.25
65.40
75.49
72.94
70.90
68.96
65.68
80.10
75.66
76.08
79.90
Similarity objective
74.39
70.30
65.50
80.97
78.30
67.90
68.90
65.81
82.80
76.06
77.04
81.50
Ours (binary mask)
81.73
79.03
58.20
89.65
88.65
45.60
70.45
66.63
76.90
77.69
77.73
82.40
PCA projection
70.33
65.08
69.80
70.52
67.91
77.50
63.07
59.18
87.80
67.89
66.36
81.20
Table 4: Ablation study on feature masking strategies and training objectives. Results are averaged across four OOD datasets (UVEB, MedVidBench, HAM10000, and WikiArt). We compare our masking strategy against the standard similarity objective ( Q1 ) and unsupervised feature-selection baselines ( Q2 ) evaluated at equal sparsity ( k=329 for V-JEPA 2.1-L, k=322 for VideoMAEv2-B).
Figure 4: PCA visualization of V-JEPA 2.1-L patch-level features mapped to RGB channels. (Top) Original input frames. (Middle) Full V-JEPA features. (Bottom) Our masked representations.
Figure 5: Quantitative information loss evaluation via K-means mosaic reconstruction. The plot compares the Mean Squared Error (MSE) between the original RGB frames and the mosaic reconstructions across varying cluster counts ( k ). The V-JEPA 2.1-L masked representations evaluate multiple binarization thresholds ( θ ).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure C.1: Per-prompt intra-set feature-space diversity for videos generated by CogVideoX1.5-5B (top row) and Mochi 1 (bottom row), measured using V-JEPA 2.1 Large (left column) and VideoMAEv2 Base (right column). Each point represents one prompt and reports the average pairwise normalized cosine dissimilarity among three videos generated with different random seeds. Specialized prompts exhibit higher intra-prompt diversity than general prompts in both feature spaces, indicating lower consistency across generations.
Figure C.2: Cross-generator qualitative comparison of intra-prompt consistency. For each general and specialized prompt, we show three videos generated with different random seeds using Wan 2.1, CogVideoX1.5, and Mochi 1. Each tile contains sampled frames from one generated video. General prompts produce comparatively consistent generations across seeds, whereas specialized prompts exhibit greater visual and semantic variation across all three generators.
Main Action
Main Subjects
Main Objects
Location
Actions Order
Average
Encoder
Method
Sp. (%)
ρ
τ
ρ
τ
ρ
τ
ρ
τ
ρ
τ
ρ
τ
Baseline (no mask)
–
24.0
16.30
33.4
22.84
24.7
16.83
49.3
35.04
24.3
16.65
31.14
21.53
Random mask
67.8
22.8
15.42
31.8
21.77
23.3
15.84
47.3
33.53
22.8
15.61
29.60
20.43
Variance (top- k )
67.8
21.7
14.67
31.4
21.51
21.3
14.51
47.8
33.75
22.1
15.16
28.86
19.92
PCA loading
67.8
21.5
14.54
31.3
21.45
21.1
14.43
47.8
33.76
21.9
15.10
28.72
19.86
Laplacian score
67.8
21.7
14.66
30.9
21.16
20.7
14.14
47.0
33.26
22.0
15.10
28.46
19.66
Appendix
Table C.1: Alignment with human judgement on ConViS-Bench. Spearman’s ρ and Kendall’s τ correlations with human judgement ( ×100 ). Sp. denotes the sparsity of the mask; soft masks are dense by construction. Unsupervised feature selectors are fit on the video embeddings alone and matched to the sparsity of our learned mask. Best results are bolded , second-best underlined .
Main Action
Main Subjects
Main Objects
Location
Actions Order
Average
Encoder
Method
Sp. (%)
ρ
τ
ρ
τ
ρ
τ
ρ
τ
ρ
τ
ρ
τ
Baseline (no mask)
–
24.0
16.30
33.4
22.84
24.7
16.83
49.3
35.04
24.3
16.65
31.14
21.53
Binary mask (CogVideoX)
75.3
30.8
21.02
37.9
26.12
31.4
21.46
51.7
36.94
34.2
23.67
37.20
25.84
Binary mask (Mochi)
72.1
28.2
19.30
37.5
25.91
30.1
20.48
52.3
37.44
30.2
20.77
35.66
24.78
Soft mask (CogVideoX)
–
31.8
21.76
38.9
26.94
32.3
22.08
53.0
37.88
34.7
24.04
38.14
26.54
V-JEPA 2.1-L
Soft mask (Mochi)
–
29.4
20.04
38.0
26.26
31.6
21.47
53.0
37.97
31.9
22.04
36.78
25.56
Appendix
Table C.2: Effect of the mask source generator on ConViS-Bench. Spearman’s ρ and Kendall’s τ correlations with human judgement ( ×100 ). Sp. denotes the sparsity of the learned mask; soft masks are dense by construction. Masks learned from synthetic datasets generated by either CogVideoX1.5-5B or Mochi 1 improve average alignment with human judgement over the corresponding unmasked representations, demonstrating robustness to the choice of source video generator. Best results are bolded , second-best underlined .
V-JEPA 2.1-L
VideoMAEv2-B
2XDM
DDPM-OOD
2XDM
DDPM-OOD
OOD Dataset
Method
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
UVEB
Unmasked
78.42
73.77
57.20
76.37
70.31
63.60
71.08
68.50
77.60
75.50
73.33
83.20
Binary mask
84.19
82.29
52.80
89.43
88.51
42.40
73.18
70.24
71.20
78.97
76.73
84.40
Soft mask
84.63
82.98
52.80
89.43
88.54
42.00
72.98
70.03
71.20
78.55
76.38
84.40
MedVidBench
Unmasked
74.66
74.18
76.00
68.11
68.47
86.00
78.59
80.73
82.40
77.01
79.46
79.60
Appendix
Table C.3: OOD detection using masks learned from CogVideoX1.5-5B. Video generation OOD detection across four datasets using Cosmos-Predict2, with BDD100k as the in-distribution (ID) baseline. The soft masks are learned exclusively from synthetic videos generated by CogVideoX1.5-5B, while the corresponding binary masks are obtained by post-hoc thresholding at θ=0.5 . The resulting improvements over the unmasked representations demonstrate transfer across source video generators and generative tasks. Best results are bolded and second-best results are underlined .
Mochi 1
CogVideoX1.5-5B
Dataset
Encoder
Unmasked
Sp. (%)
Binary
Soft
Sp. (%)
Binary
Soft
UCF101
V-JEPA 2.1 Large
40.82
72.2
42.84
43.48
75.4
44.02
44.42
VideoMAEv2-Base
95.52
52.0
95.56
95.62
57.8
95.55
95.61
HMDB51
V-JEPA 2.1 Large
17.84
72.2
18.99
19.59
75.4
19.91
20.19
VideoMAEv2-Base
46.94
52.0
46.59
46.88
57.8
46.51
46.74
Appendix
Table C.4: Class-level video retrieval mAP (%) on UCF101 and HMDB51, without and with the masks learned from Mochi 1 and CogVideoX generated datasets. Sp. denotes mask sparsity in %. Best results are bolded and second-best results are underlined per generator.
V-JEPA2.1-L
VideoMAEv2-B
2XDM
DDPM-OOD
2XDM
DDPM-OOD
OOD Dataset
Method
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
UVEB
Unmasked
78.42
73.77
57.20
76.37
70.31
63.60
71.08
68.50
77.60
75.50
73.33
83.20
Binary mask
81.84
78.98
54.00
84.85
81.37
51.20
72.32
69.96
75.20
78.01
76.01
83.20
Soft mask
82.61
80.33
52.80
85.99
83.20
48.80
72.19
69.92
74.80
77.66
75.76
83.60
MedVidBench
Unmasked
74.66
74.18
76.00
68.11
68.47
86.00
78.59
80.73
82.40
77.01
79.46
79.60
Appendix
Table C.5: OOD detection using masks learned from Mochi 1. Video generation OOD detection across four datasets using Cosmos-Predict2, with BDD100k as the in-distribution (ID) baseline. The soft masks are learned exclusively from synthetic videos generated by Mochi 1, while the corresponding binary masks are obtained by post-hoc thresholding at θ=0.5 . The improvements across diverse OOD domains provide additional evidence that the proposed learning signal transfers across source video generators. Best results are bolded and second-best results are underlined .
Figure C.3: Patch-level effectiveness of masked features. VOS on DAVIS-2017 ( Pont-Tuset et al., 2017 ) via patch features matching.
Figure C.4: Qualitative visualization of K-means mosaic reconstructions. The original input frame (left) is segmented based on patch-level cluster assignments for k∈{2,4,8,16,32,64} . The visual consistency between the unmasked baseline (top row) and our masked representations (bottom row) illustrates that the masked feature space preserves essential local structural and color information.
Despite remarkable advances in video generative models, they still struggle to generate physically realistic videos, frequently exhibiting appearance drift, implausible motion, and temporal inconsistencies. In this work, we address this limitation by transferring relational knowledge encoded in spatio-temporal self-similarity (STSS) from visual foundation models into video generative models. STSS represents pairwise similarities among features across space and time, revealing the relational structure of how objects interact with other entities throughout a video, effectively capturing real-world dynamics, including object motion and semantic transformations. To transfer this relational knowledge, we propose Tempered Self-similarity Alignment (TSA) loss, which transforms STSS into probabilistic correspondence distributions and trains the video generative model to align its correspondence distributions with those of the visual foundation model on dynamically changing regions. Evaluated on VideoPhy and VideoPhy2 benchmarks, our method demonstrates substantial improvements in physical plausibility across diverse interaction scenarios, validating the effectiveness of transferring relational knowledge for physically realistic video generation.
Manjin Kim, Suha Kwak, Minsu Cho
Pohang University of Science and Technology (POSTECH)
Video generation models have demonstrated emerging zero-shot capabilities for visual reasoning, perception, and other vision tasks. However, diffusion-based video generation is inherently stochastic, while many downstream vision tasks are deterministic. Motivated by the effectiveness of self-consistency in chain-of-thought reasoning for large language models, we investigate whether self-consistency can similarly improve diffusion-based video reasoning. We first introduce a training-free test-time scaling method that samples multiple video generations and aggregates their predictions through self-consistency. Specifically, we aggregate extracted paths, locations, or masks from multiple rollouts into a consensus prediction. To reduce the inference overhead of multi-rollout generation, we read out predictions early in the denoising trajectory, which preserves consensus quality while reducing denoising steps by more than half. We further propose Rejection Fine-Tuning (RFT) to distill consensus predictions into the video generation model. The resulting model internalizes the benefit of multi-sample consensus and requires only a single generation at inference time, while substantially outperforming the original model. Experiments on three tasks, including maze solving, visual search, and referring segmentation, show that both our self-consistency inference and consensus distillation dramatically improve video-based perception and reasoning, without requiring ground-truth videos or task-specific verification. For visual search, self-consistency raises task accuracy from 48.4% for a single generation to 99.0%. The distilled model retains much of the consensus benefit with a single rollout. For 4-by-4 maze solving, consensus-based training improves the single-generation strict success rate from 72.0% to 84.0% with the same inference latency.
Zhenghao Ni, Weimin Qiu, Meng Tang
University of Toronto · University of California, Merced
Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models overcome this limitation. By systematically probing their intermediate activations using recent mutual-kNN alignment metrics, we reveal a highly structured latent space where visual representations evolve across both network depth and noise levels. We show that while moderate noise levels yield linearly separable global semantics, fine-grained details persist at lower noise levels but become spatially scattered, requiring attention mechanisms to decode. Building on these insights, we introduce Gen4U (Generation for Understanding), a framework that repurposes these generative representations with a single forward pass. Our experiments establish that frozen, large-scale video diffusion models function as highly competitive video encoders across a wide spectrum of tasks, spanning semantic and non-semantic objectives (video classification, depth estimation, camera pose estimation, image and video captioning). Bypassing fine-tuning, Gen4U unifies the generation and understanding paradigms, achieving strong perception performance while fully preserving the model's ability to generate high-quality video.
Michael King, Aravindh Mahendran, Matthew Koichi Grimes +5