cs.LGSep 23, 2026

LAYERSCOPE: A Layerwise Characterization of Video and Multimodal Learned Representations

Authors: Sandra Arcos-Holzinger, Debashish Chakraborty, Rohita Mocharla, Will Walden, Andrew Yates, Reno Kriz, Sarah M. Erfani, James Bailey, +2 more

Organizations: Johns Hopkins University · University of Melbourne · Human Language Technology Center of Excellence · Johns Hopkins Applied Physics Lab · Monash University

Abstract

We propose LAYERSCOPE, a label-free, layerwise framework that aims to characterize a model's learned representations in video and multimodal settings. Evaluating downstream performance using representations from final or intermediate layers typically requires large amounts of labeled data, repeated task-specific evaluations, and substantial computation. To address these limitations, LAYERSCOPE uses local, global, distributional, and correspondence-based geometric metrics to compare layerwise representation structure within and across models without requiring task-specific labels. We evaluate seven architecturally diverse models across video and multimodal classification, clustering, and text-to-video retrieval tasks from MVEB/MVEB+. We find that intermediate-layer representations can outperform final-layer and model-default outputs. We also find that no single geometric metric consistently predicts downstream performance, but note that distinct layerwise geometric signatures emerge across model families. LID shows task-dependent relationships with performance, while RankMe provides the strongest measure for classification and clustering, but is not a universal layer selector. We also find that pairing-aware metrics explain retrieval better than distributional distances alone. LAYERSCOPE therefore offers a framework for comparing representations across models and layers, enabling a more systematic evaluation in video and multimodal settings.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Uncovering the Latent Potential of Deep Intermediate Representations

    May 21, 2026Arnesh Batra, Arush Gumber, Aniket Khandelwal +2Layer-WiseCross-Modal

  2. How Far Are Video Models from True Multimodal Reasoning?

    Apr 21, 2026Xiaotian Zhang, Jianhui Wei, Yuan Wang +9Multimodal ReasoningVideo Understanding

  3. HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding

    May 19, 2026Mengqi Shi, Haopeng ZhangMultimodal UnderstandingMultimodal Benchmarks