LAYERSCOPE: A Layerwise Characterization of Video and Multimodal Learned Representations
Authors: Sandra Arcos-Holzinger, Debashish Chakraborty, Rohita Mocharla, Will Walden, Andrew Yates, Reno Kriz, Sarah M. Erfani, James Bailey, +2 more
Organizations: Johns Hopkins University · University of Melbourne · Human Language Technology Center of Excellence · Johns Hopkins Applied Physics Lab · Monash University
We propose LAYERSCOPE, a label-free, layerwise framework that aims to characterize a model's learned representations in video and multimodal settings. Evaluating downstream performance using representations from final or intermediate layers typically requires large amounts of labeled data, repeated task-specific evaluations, and substantial computation. To address these limitations, LAYERSCOPE uses local, global, distributional, and correspondence-based geometric metrics to compare layerwise representation structure within and across models without requiring task-specific labels. We evaluate seven architecturally diverse models across video and multimodal classification, clustering, and text-to-video retrieval tasks from MVEB/MVEB+. We find that intermediate-layer representations can outperform final-layer and model-default outputs. We also find that no single geometric metric consistently predicts downstream performance, but note that distinct layerwise geometric signatures emerge across model families. LID shows task-dependent relationships with performance, while RankMe provides the strongest measure for classification and clustering, but is not a universal layer selector. We also find that pairing-aware metrics explain retrieval better than distributional distances alone. LAYERSCOPE therefore offers a framework for comparing representations across models and layers, enabling a more systematic evaluation in video and multimodal settings.
Figures & tables
Figure 1: LayerScope framework. Given video-only tasks or tasks evaluated in a shared text–video space, LayerScope freezes each model and evaluates representations across transformer depth. For classification and clustering, eligible post-block hidden states are mean pooled; model-default outputs are evaluated separately. For models with shared text–video spaces, the model-specific normalization, pooling, and projection operations are applied at every text and video layer, producing the complete pair matrix p=(ℓT,ℓV)∈Pm,t . LayerScope relates GRIDS-LID, RankMe, covariance effective rank, sliced 2-Wasserstein distance, and held-out orthogonal Procrustes error to downstream performance through within-model association and geometry-only selection regret.
Figure 2: Full-pool UCF101 geometry signatures grouped by architecture and readout family. Spectral values are normalized for display only; accepted associations and selectors use their absolute values.
Figure 3: Layer-choice regret by task and model. Each cell shows the regret from selecting a layer using a geometric measure. Lower values are better; this makes it clear that the usefulness of a geometric measure depends on the task and model.
Figure 4: UCF101 layerwise profiles for seven models. GRIDS-LID, classification and clustering means are plotted layerwise against normalized model depth. For each model profile, diamonds in each shaded margin mark the model-default output. Task scores are on a 0–1 scale.
Figure 5: Cross-modal distributional alignment and retrieval performance across normalized text and vision depth. The top row shows average Sliced Wasserstein distance (SW 2 ) between text and vision representations, while the bottom rows show text-to-video retrieval performance (nDCG@10) on TUNA and VATEX. Although lower SW 2 often coincides with improved retrieval, we find that cross-modal layer pairs, with minimum SW 2 , do not consistently result in a high retrieval performance.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Model Family
Model
Checkpoint
Modalities
Architecture
Predictive video
V-JEPA 2
facebook/vjepa2-vitl-fpc64-256
Video
ViT-L
Audio-video contrastive
PE-AV Small
facebook/pe-av-small
Text/video
Video + text encoders
Audio-video contrastive
PE-AV Large
facebook/pe-av-large
Text/video
Video + text encoders
Video-text contrastive
X-CLIP
microsoft/xclip-base-patch32
Text/video
Base dual encoder
Autoregressive multimodal
Gemma 4
google/gemma-4-E4B-it
Multimodal
E4B-IT decoder
Autoregressive multimodal
Qwen3-VL
Qwen/Qwen3-VL-Embedding-2B
Text/video
Embedding-2B decoder
Appendix
Table 1: Models and checkpoints used in the experiments. Modalities indicate the inputs used in this study.
Analysis
Dataset
Evaluated representations
Evaluation
Score
Classification
Breakfast
Every video layer; model-default output
Eight-shot linear probe
Accuracy
Classification
UCF101
Every video layer; model-default output
Eight-shot linear probe
Accuracy
Classification
HMDB51
Every video layer; model-default output
Eight-shot linear probe
Accuracy
Clustering
UCF101
Every video layer; model-default output
MiniBatchKMeans
V-measure
Text-to-video retrieval
TUNA-Bench
Every text–video layer pair; model-default pair
Cosine ranking
Mean nDCG@10
Text-to-video retrieval
VATEX
Every text–video layer pair; model-default pair
Cosine ranking
Mean nDCG@10
Appendix
Table 3: Downstream evaluations. Classification and clustering scores are summarized over ten evaluator seeds. Retrieval and zero-shot classification are deterministic for fixed representations.
Figure 6: Sensitivity of GRIDS-LID to neighborhood size. Panel A compares the k=50 and k=150 trajectories with the primary k=100 trajectory over 36 representation views and reports agreement of the minimum- and maximum-GRIDS-LID layers. Panel B reports layer-choice regret over the 21 primary classification and clustering comparisons, using evaluator seeds 0 – 9 . Regret is shown on the 0 – 1 downstream-score scale.
Figure 7: SW2 and nDCG@10 across all 5218 accepted layer-pair cells. Cyan squares mark minimum SW 2 , red stars the retrieval oracle, and black diamonds the detached default pair. Lower marginal distance often tracks retrieval but does not identify the optimum; axes are model-specific.
Task
Model
Matrix
Oracle (ℓT,ℓV)
Oracle score
Matched depth
Δmatched
Model-default
Δdefault
TUNA
PE-AV Small
22×4
19/4
0.9681
0.9681
+0.0000
0.9671
+0.0010
VATEX
PE-AV Small
22×4
21/4
0.7914
0.7914
+0.0000
0.7899
+0.0016
TUNA
PE-AV Large
22×4
22/4
0.9695
0.9695
+0.0000
0.9695
+0.0000
VATEX
PE-AV Large
22×4
22/4
0.8172
0.8172
+0.0000
0.8172
+0.0000
TUNA
X-CLIP
12×12
12/11
0.4014
0.3329
+0.0685
0.3328
+0.0686
VATEX
X-CLIP
12×12
12/11
0.4148
0.3950
+0.0198
0.3950
+0.0198
Appendix
Table 4: Text-to-video retrieval over all text–video layer pairs. The oracle is the pair with the highest nDCG@10. The matched-depth optimum is selected from the normalized-depth path, while the model-default output is evaluated separately. Scores and differences use the 0 – 1 nDCG@10 scale.
Figure 8: Linear CKA and held-out orthogonal Procrustes error for text–video layer pairs aligned by relative depth on TUNA and VATEX. Higher CKA and lower Procrustes error are favorable. The figure includes all five shared-space models and uses every evaluated text layer.
Figure 9: Separation between the observed text–video pairing and 1,000 correspondence permutations across 32 model–task–depth comparisons.
Model
Best (ℓT,ℓV)
Best accuracy
Model-default
Difference
PE-AV Small
18/4
0.7449
0.7366
+0.0082
PE-AV Large
21/4
0.8421
0.8390
+0.0031
X-CLIP
12/12
0.7598
0.7603
−0.0005
Qwen3-VL
27/27
0.8724
0.8719
+0.0005
LCO
35/35
0.8323
0.8302
+0.0021
Appendix
Table 5: UCF101 zero-shot classification with four-prompt ensemble. Here, the best pair is selected from all text–video layer pairs.
Model
Task
Layers
ℓbest
sbest
sfinal
Δfinal
sdefault
Δdefault
V-JEPA 2
Breakfast cls.
24
22
0.3325±0.0243
0.3110
+0.0215
0.3124
+0.0201
V-JEPA 2
UCF101 cls.
24
23
0.8129±0.0045
0.8072
+0.0058
0.8062
+0.0068
V-JEPA 2
UCF101 clust.
24
23
0.7707±0.0040
0.7702
+0.0006
0.7830
−0.0123
PE-AV Small
Breakfast cls.
4
4
0.5185±0.0228
0.5185
+0.0000
0.4529
+0.0656
PE-AV Small
UCF101 cls.
4
1
0.9338±0.0027
0.9248
+0.0090
0.8759
+0.0579
PE-AV Small
UCF101 clust.
4
1
0.8283±0.0040
0.7589
+0.0694
0.8508
−0.0225
Appendix
Table 6: Layerwise classification (cls.) and clustering (clust.) results on Breakfast and UCF101. Classification is measured by accuracy and clustering by V-measure. ℓbest is the layer with the highest mean score across ten evaluator seeds. sbest , sfinal , and sdefault denote the corresponding ten-seed mean scores; the standard deviation shown with sbest is computed at the fixed layer ℓbest .
Model
ℓbest
sbest
sfinal
sdefault
Δfinal
Δdefault
V-JEPA 2
23
0.4690
0.4682
0.4578
+0.0009
+0.0113
PE-AV Small
1
0.6284
0.5987
0.4247
+0.0297
+0.2036
Gemma 4
14
0.4724
0.3904
0.0891
+0.0820
+0.3833
PE-AV Large
2
0.6462
0.6237
0.5136
+0.0225
+0.1326
X-CLIP
12
0.4807
0.4807
0.6242
+0.0000
−0.1435
Qwen3-VL
28
0.6224
0.6224
0.6602
+0.0000
−0.0378
Appendix
Table 7: HMDB51 eight-shot classification across all evaluated video layers. ℓbest is selected by mean accuracy across ten evaluator seeds. sbest , sfinal , and sdefault are the corresponding mean accuracies. Positive differences favor the best raw layer.
Figure 10: Within-model, within-task Spearman associations between GRIDS-LID, RankMe, standardized covariance effective rank, and downstream performance across layers. Task scores are ten-seed means. lr denotes a four-layer trajectory, for which we do not report a primary correlation; ud denotes an undefined association. Correlations use the [−1,1] scale.
Layer choice
N
Mean
Median
Maximum
Zero
W/T/L vs. last
Maximum RankMe
21
0.0143
0.0081
0.0558
3
8/9/4
Last layer
21
0.0249
0.0090
0.1694
4
–
Maximum standardized covariance effective rank
21
0.0399
0.0147
0.2172
3
10/3/8
Minimum GRIDS-LID ( k=100 )
21
0.0447
0.0063
0.2380
4
7/4/10
Maximum GRIDS-LID ( k=100 )
21
0.0997
0.0374
0.4734
0
3/6/12
Appendix
Table 8: Layer-choice regret across the 21 primary classification and clustering comparisons. Regret uses the native 0 – 1 task-score scale; lower values are better. “Zero” counts comparisons with zero mean regret. Win/tie/loss compares each geometric choice with the last-layer baseline.
Figure 11: Layer-choice regret for each of the 21 primary classification and clustering comparisons. Each cell reports the mean and sample standard deviation over evaluator seeds 0 – 9 on the 0 – 1 task-score scale. Minimum and maximum GRIDS-LID are evaluated separately, and the last layer is a fixed baseline. Asterisks mark values above the displayed 95th-percentile color cap.
Foundational Models pretrained on huge amount of data learn representations that evolve across depth, forming a hierarchy of embeddings with distinct semantic content and geometric structure. Contrary to the widespread practice of using only the final layer or shallow mixtures, we show that task-relevant information is distributed non-monotonically across layers and cannot be recovered by naïve aggregation. Through a geometric and empirical study across multiple modalities, we show that effective transfer depends on identifying which layers encode task-discriminative structure and how their embeddings are geometrically organized. We introduce Layer-wise Optimal Embedding Selection (LOES), a constructive spectral method that identifies task-discriminative subspaces by minimizing residual error under orthogonality and isotropy constraints. To align fine-tuning with this selection principle, we further propose Geometric Regularization Loss (GeoReg), which enforces a simplicial structure on class manifolds and stabilizes representation geometry during fine-tuning. Across a wide range of architectures, depths, modalities, and data regimes, LOES consistently outperforms standard baselines, with gains that grow as model depth increases. Beyond accuracy, our method reveals how semantic factors are distributed across layers, thereby enabling cross-lingual and cross-modal interpretability analyses. Together, our results provide strong evidence that layerwise embedding geometry is not incidental but central to how deep models represent and transfer knowledge.
Arnesh Batra, Arush Gumber, Aniket Khandelwal +2
SBILab, Indraprastha Institute of Information Technology Delhi, Delhi, India.
Despite remarkable progress toward general-purpose video models, a critical question remains unanswered: how far are these models from achieving true multimodal reasoning? Existing benchmarks fail to address this question rigorously, as they remain constrained by straightforward task designs and fragmented evaluation metrics that neglect complex multimodal reasoning. To bridge this gap, we introduce CLVG-Bench, an evaluation framework designed to probe video models' zero-shot reasoning capabilities via Context Learning in Video Generation. CLVG-Bench comprises more than 1,000 high-quality, manually annotated metadata across 6 categories and 47 subcategories, covering complex scenarios including physical simulation, logical reasoning, and interactive contexts. To enable rigorous and scalable assessment, we further propose an Adaptive Video Evaluator (AVE) that aligns with human expert perception using minimal annotations, delivering interpretable textual feedback across diverse video context tasks. Extensive experiments reveal a striking answer to our central question: while state-of-the-art (SOTA) video models, such as Seedance 2.0, demonstrate competence on certain understanding and reasoning subtasks, they fall substantially short with logically grounded and interactive generation tasks (achieving success rates <25% and ~0%, respectively), exposing multimodal reasoning and physical grounding as critical bottlenecks. By systematically quantifying these limitations, the proposed method provides actionable feedbacks and a clear roadmap toward truly robust, general-purpose video models. CLVG-Bench and code are released here.
While Multimodal Large Language Models (MLLMs) exhibit strong performance on standard video tasks, their ability to faithfully summarize and reason over complex narratives remains poorly evaluated. Existing summarization benchmarks fragment supervision across isolated granularities, such as keyframes, key shots, or disjointed text summaries, failing to capture the inherently hierarchical structure of cross-modal alignment. To address this critical gap, we introduce HAVEN, a hierarchically aligned multimodal benchmark for unified video understanding. HAVEN pioneers a fully granular (frame, shot, and video levels) and fully multimodal (video and text) dataset architecture, complete with explicit, continuous alignment between modalities. Built upon this unified annotation paradigm, we propose a comprehensive evaluation suite spanning summarization, temporal reasoning, multimodal grounding, and saliency ranking. Extensive benchmarking of state-of-the-art MLLMs exposes a persistent gap between surface-level textual fluency and grounded multimodal understanding. Ultimately, HAVEN advances the evaluation of multimodal systems beyond traditional QA formats, offering a rigorous, standardized testbed to drive future research in interpretable, hierarchical video understanding. We publicly release the dataset, benchmark suite, and evaluation protocols.
Mengqi Shi, Haopeng Zhang
Department of Information and Computer Sciences · University of Hawaii at Manoa