Evaluating semantic similarity between videos is a fundamental challenge in computer vision, essential for tasks ranging from out-of-distribution (OOD) detection to video retrieval. However, defining and labeling video similarity is notoriously difficult and expensive due to the complex spatio-temporal nature. In this paper, we propose a novel self-supervised approach that leverages generative uncertainty from text-to-video (T2V) diffusion models to learn semantic similarity without human annotations. Our method is based on the observation that T2V models produce consistent outputs for familiar concepts but exhibit high variance and uncertainty when prompted with specialized concepts. We utilize this behavior to identify stable semantic features within existing pretrained representations, such as VideoMAE and V-JEPA. Specifically, we learn a mask over these embeddings using purely generated data, encouraging the model to retain features that remain consistent across generations of general concepts while discarding those associated with generative noise or uncertainty. Experimental results across three key tasks demonstrate that our learned feature subspaces consistently outperform original pretrained features and baseline feature selection methods.
Figures & tables
Figure 1: Proposed framework overview. (1) A text-to-video model generates videos from general and specialized prompts with different random seeds; videos from the same prompt are paired. (2) A frozen video encoder extracts global representations, and a learnable mask is trained to yield high/low cosine similarity for similar/dissimilar pairs. (3) For downstream tasks, we mask unseen video representations before cosine similarity. In doing so, our strategy improves alignment with human judgment, out-of-distribution detection, and video retrieval.
Figure 2: Example videos used in our generative consistency objective. Videos generated by Wan 2.1 with general and specialized text prompts. Videos from general prompts serve as proxies for semantic consistency, while those from specialized prompts represent generative divergence.
Figure 3: Per-prompt intra-set feature-space diversity for videos generated by Wan2.1, measured using V-JEPA 2.1 Large (left) and VideoMAEv2 Base (right). Each point represents one prompt and reports the average pairwise normalized cosine dissimilarity among three videos generated with different random seeds.
Main Action
Main Subjects
Main Objects
Location
Actions Order
Average
Encoder
Method
Sp. (%)
ρ
τ
ρ
τ
ρ
τ
ρ
τ
ρ
τ
ρ
τ
Baseline (no mask)
–
24.0
16.30
33.4
22.84
24.7
16.83
49.3
35.04
24.3
16.65
31.14
21.53
Binary mask
67.8
26.9
18.25
39.4
27.14
30.3
20.57
50.5
35.74
30.1
20.75
35.44
24.49
V-JEPA 2.1-L
Soft mask
–
27.7
18.78
39.5
27.24
30.8
20.89
51.3
36.41
30.9
21.23
36.04
24.91
Baseline (no mask)
–
47.0
33.02
44.6
31.46
42.4
29.96
45.8
32.26
44.7
31.37
44.90
31.61
Binary mask
58.0
47.9
33.62
46.3
32.63
44.1
30.98
46.9
33.09
45.3
31.67
46.10
32.40
Table 1: Alignment with human judgement on ConViS-Bench. Spearman’s ρ and Kendall’s τ correlations with human judgement ( ×100 ). Sp. denotes the sparsity of the learned mask; soft masks are dense by construction. Our masked features better match the human notion of global video similarity. Best results are bolded , second-best underlined .
V-JEPA 2.1-L
VideoMAEv2-B
2XDM
DDPM-OOD
2XDM
DDPM-OOD
OOD Dataset
Method
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
UVEB
Unmasked
78.42
73.77
57.20
76.37
70.31
63.60
71.08
68.50
77.60
75.50
73.33
83.20
Binary mask
86.77
85.72
51.20
89.52
88.04
42.80
72.31
69.27
73.60
77.47
75.10
84.80
Soft mask
85.79
84.48
52.00
89.11
87.35
42.80
72.21
68.90
73.20
77.37
74.81
83.20
MedVidBench
Unmasked
74.66
74.18
76.00
68.11
68.47
86.00
78.59
80.73
82.40
77.01
79.46
79.60
Table 2: Video generation OOD detection across four diverse datasets. Videos are generated using Cosmos-Predict2, with BDD100k as the in-distribution (ID) baseline. Applying binary and soft feature masks consistently improves OOD detection performance. Best results are bolded and second-best results are underlined .
Dataset
Encoder
Retrieval mAP (%)
Unmasked
Binary mask
Soft mask
UCF101
V-JEPA 2.1 Large
40.82
41.34
42.37
VideoMAEv2-Base
95.52
95.61
95.65
HMDB51
V-JEPA 2.1 Large
17.84
18.40
18.99
VideoMAEv2-Base
46.94
46.39
46.66
Table 3: Class-level video retrieval mAP (%) on UCF101 and HMDB51. Best results are bolded and second-best results are underlined .
V-JEPA 2.1-L
VideoMAEv2-B
2XDM
DDPM-OOD
2XDM
DDPM-OOD
Method
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
Baseline (unmasked)
71.75
66.25
65.40
75.49
72.94
70.90
68.96
65.68
80.10
75.66
76.08
79.90
Similarity objective
74.39
70.30
65.50
80.97
78.30
67.90
68.90
65.81
82.80
76.06
77.04
81.50
Ours (binary mask)
81.73
79.03
58.20
89.65
88.65
45.60
70.45
66.63
76.90
77.69
77.73
82.40
PCA projection
70.33
65.08
69.80
70.52
67.91
77.50
63.07
59.18
87.80
67.89
66.36
81.20
Table 4: Ablation study on feature masking strategies and training objectives. Results are averaged across four OOD datasets (UVEB, MedVidBench, HAM10000, and WikiArt). We compare our masking strategy against the standard similarity objective ( Q1 ) and unsupervised feature-selection baselines ( Q2 ) evaluated at equal sparsity ( k=329 for V-JEPA 2.1-L, k=322 for VideoMAEv2-B).
Figure 4: PCA visualization of V-JEPA 2.1-L patch-level features mapped to RGB channels. (Top) Original input frames. (Middle) Full V-JEPA features. (Bottom) Our masked representations.
Figure 5: Quantitative information loss evaluation via K-means mosaic reconstruction. The plot compares the Mean Squared Error (MSE) between the original RGB frames and the mosaic reconstructions across varying cluster counts ( k ). The V-JEPA 2.1-L masked representations evaluate multiple binarization thresholds ( θ ).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure C.1: Per-prompt intra-set feature-space diversity for videos generated by CogVideoX1.5-5B (top row) and Mochi 1 (bottom row), measured using V-JEPA 2.1 Large (left column) and VideoMAEv2 Base (right column). Each point represents one prompt and reports the average pairwise normalized cosine dissimilarity among three videos generated with different random seeds. Specialized prompts exhibit higher intra-prompt diversity than general prompts in both feature spaces, indicating lower consistency across generations.
Figure C.2: Cross-generator qualitative comparison of intra-prompt consistency. For each general and specialized prompt, we show three videos generated with different random seeds using Wan 2.1, CogVideoX1.5, and Mochi 1. Each tile contains sampled frames from one generated video. General prompts produce comparatively consistent generations across seeds, whereas specialized prompts exhibit greater visual and semantic variation across all three generators.
Main Action
Main Subjects
Main Objects
Location
Actions Order
Average
Encoder
Method
Sp. (%)
ρ
τ
ρ
τ
ρ
τ
ρ
τ
ρ
τ
ρ
τ
Baseline (no mask)
–
24.0
16.30
33.4
22.84
24.7
16.83
49.3
35.04
24.3
16.65
31.14
21.53
Random mask
67.8
22.8
15.42
31.8
21.77
23.3
15.84
47.3
33.53
22.8
15.61
29.60
20.43
Variance (top- k )
67.8
21.7
14.67
31.4
21.51
21.3
14.51
47.8
33.75
22.1
15.16
28.86
19.92
PCA loading
67.8
21.5
14.54
31.3
21.45
21.1
14.43
47.8
33.76
21.9
15.10
28.72
19.86
Laplacian score
67.8
21.7
14.66
30.9
21.16
20.7
14.14
47.0
33.26
22.0
15.10
28.46
19.66
Appendix
Table C.1: Alignment with human judgement on ConViS-Bench. Spearman’s ρ and Kendall’s τ correlations with human judgement ( ×100 ). Sp. denotes the sparsity of the mask; soft masks are dense by construction. Unsupervised feature selectors are fit on the video embeddings alone and matched to the sparsity of our learned mask. Best results are bolded , second-best underlined .
Main Action
Main Subjects
Main Objects
Location
Actions Order
Average
Encoder
Method
Sp. (%)
ρ
τ
ρ
τ
ρ
τ
ρ
τ
ρ
τ
ρ
τ
Baseline (no mask)
–
24.0
16.30
33.4
22.84
24.7
16.83
49.3
35.04
24.3
16.65
31.14
21.53
Binary mask (CogVideoX)
75.3
30.8
21.02
37.9
26.12
31.4
21.46
51.7
36.94
34.2
23.67
37.20
25.84
Binary mask (Mochi)
72.1
28.2
19.30
37.5
25.91
30.1
20.48
52.3
37.44
30.2
20.77
35.66
24.78
Soft mask (CogVideoX)
–
31.8
21.76
38.9
26.94
32.3
22.08
53.0
37.88
34.7
24.04
38.14
26.54
V-JEPA 2.1-L
Soft mask (Mochi)
–
29.4
20.04
38.0
26.26
31.6
21.47
53.0
37.97
31.9
22.04
36.78
25.56
Appendix
Table C.2: Effect of the mask source generator on ConViS-Bench. Spearman’s ρ and Kendall’s τ correlations with human judgement ( ×100 ). Sp. denotes the sparsity of the learned mask; soft masks are dense by construction. Masks learned from synthetic datasets generated by either CogVideoX1.5-5B or Mochi 1 improve average alignment with human judgement over the corresponding unmasked representations, demonstrating robustness to the choice of source video generator. Best results are bolded , second-best underlined .
V-JEPA 2.1-L
VideoMAEv2-B
2XDM
DDPM-OOD
2XDM
DDPM-OOD
OOD Dataset
Method
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
UVEB
Unmasked
78.42
73.77
57.20
76.37
70.31
63.60
71.08
68.50
77.60
75.50
73.33
83.20
Binary mask
84.19
82.29
52.80
89.43
88.51
42.40
73.18
70.24
71.20
78.97
76.73
84.40
Soft mask
84.63
82.98
52.80
89.43
88.54
42.00
72.98
70.03
71.20
78.55
76.38
84.40
MedVidBench
Unmasked
74.66
74.18
76.00
68.11
68.47
86.00
78.59
80.73
82.40
77.01
79.46
79.60
Appendix
Table C.3: OOD detection using masks learned from CogVideoX1.5-5B. Video generation OOD detection across four datasets using Cosmos-Predict2, with BDD100k as the in-distribution (ID) baseline. The soft masks are learned exclusively from synthetic videos generated by CogVideoX1.5-5B, while the corresponding binary masks are obtained by post-hoc thresholding at θ=0.5 . The resulting improvements over the unmasked representations demonstrate transfer across source video generators and generative tasks. Best results are bolded and second-best results are underlined .
Mochi 1
CogVideoX1.5-5B
Dataset
Encoder
Unmasked
Sp. (%)
Binary
Soft
Sp. (%)
Binary
Soft
UCF101
V-JEPA 2.1 Large
40.82
72.2
42.84
43.48
75.4
44.02
44.42
VideoMAEv2-Base
95.52
52.0
95.56
95.62
57.8
95.55
95.61
HMDB51
V-JEPA 2.1 Large
17.84
72.2
18.99
19.59
75.4
19.91
20.19
VideoMAEv2-Base
46.94
52.0
46.59
46.88
57.8
46.51
46.74
Appendix
Table C.4: Class-level video retrieval mAP (%) on UCF101 and HMDB51, without and with the masks learned from Mochi 1 and CogVideoX generated datasets. Sp. denotes mask sparsity in %. Best results are bolded and second-best results are underlined per generator.
V-JEPA2.1-L
VideoMAEv2-B
2XDM
DDPM-OOD
2XDM
DDPM-OOD
OOD Dataset
Method
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
AUROC ↑
AUPR ↑
FPR@95 ↓
UVEB
Unmasked
78.42
73.77
57.20
76.37
70.31
63.60
71.08
68.50
77.60
75.50
73.33
83.20
Binary mask
81.84
78.98
54.00
84.85
81.37
51.20
72.32
69.96
75.20
78.01
76.01
83.20
Soft mask
82.61
80.33
52.80
85.99
83.20
48.80
72.19
69.92
74.80
77.66
75.76
83.60
MedVidBench
Unmasked
74.66
74.18
76.00
68.11
68.47
86.00
78.59
80.73
82.40
77.01
79.46
79.60
Appendix
Table C.5: OOD detection using masks learned from Mochi 1. Video generation OOD detection across four datasets using Cosmos-Predict2, with BDD100k as the in-distribution (ID) baseline. The soft masks are learned exclusively from synthetic videos generated by Mochi 1, while the corresponding binary masks are obtained by post-hoc thresholding at θ=0.5 . The improvements across diverse OOD domains provide additional evidence that the proposed learning signal transfers across source video generators. Best results are bolded and second-best results are underlined .
Figure C.3: Patch-level effectiveness of masked features. VOS on DAVIS-2017 ( Pont-Tuset et al., 2017 ) via patch features matching.
Figure C.4: Qualitative visualization of K-means mosaic reconstructions. The original input frame (left) is segmented based on patch-level cluster assignments for k∈{2,4,8,16,32,64} . The visual consistency between the unmasked baseline (top row) and our masked representations (bottom row) illustrates that the masked feature space preserves essential local structural and color information.