cs.CVSep 28, 2026

Generative Uncertainty as a Self-supervised Signal for Semantic Similarity Learning

Authors: Enrico Pallotta, Sina Raoufi, Lars Doorenbos, Gianni Franchi, Juergen Gall

Organizations: University of Bonn · Lamarr Institute for ML & AI · ENSTA Paris, Institut Polytechnique de Paris

Abstract

Evaluating semantic similarity between videos is a fundamental challenge in computer vision, essential for tasks ranging from out-of-distribution (OOD) detection to video retrieval. However, defining and labeling video similarity is notoriously difficult and expensive due to the complex spatio-temporal nature. In this paper, we propose a novel self-supervised approach that leverages generative uncertainty from text-to-video (T2V) diffusion models to learn semantic similarity without human annotations. Our method is based on the observation that T2V models produce consistent outputs for familiar concepts but exhibit high variance and uncertainty when prompted with specialized concepts. We utilize this behavior to identify stable semantic features within existing pretrained representations, such as VideoMAE and V-JEPA. Specifically, we learn a mask over these embeddings using purely generated data, encouraging the model to retain features that remain consistent across generations of general concepts while discarding those associated with generative noise or uncertainty. Experimental results across three key tasks demonstrate that our learned feature subspaces consistently outperform original pretrained features and baseline feature selection methods.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Tempered Self-Similarity Alignment for Physically Plausible Video Generation

    May 24, 2026Manjin Kim, Suha Kwak, Minsu ChoGenerative Video ModelsGenerative Models

  2. Learning via Self-Consistency for Diffusion-based Video Reasoning

    Sep 29, 2026Zhenghao Ni, Weimin Qiu, Meng TangVideo GenerationSelf-Consistency

  3. Gen4U: Unifying Video Generation and Understanding via Diffusion

    Jul 7, 2026Michael King, Aravindh Mahendran, Matthew Koichi Grimes +5Video Diffusion ModelsVideo Generation