LAYERSCOPE: A Layerwise Characterization of Video and Multimodal Learned Representations
Organizations: Johns Hopkins University · University of Melbourne · Human Language Technology Center of Excellence · Johns Hopkins Applied Physics Lab · Monash University
Abstract
We propose LAYERSCOPE, a label-free, layerwise framework that aims to characterize a model's learned representations in video and multimodal settings. Evaluating downstream performance using representations from final or intermediate layers typically requires large amounts of labeled data, repeated task-specific evaluations, and substantial computation. To address these limitations, LAYERSCOPE uses local, global, distributional, and correspondence-based geometric metrics to compare layerwise representation structure within and across models without requiring task-specific labels. We evaluate seven architecturally diverse models across video and multimodal classification, clustering, and text-to-video retrieval tasks from MVEB/MVEB+. We find that intermediate-layer representations can outperform final-layer and model-default outputs. We also find that no single geometric metric consistently predicts downstream performance, but note that distinct layerwise geometric signatures emerge across model families. LID shows task-dependent relationships with performance, while RankMe provides the strongest measure for classification and clustering, but is not a universal layer selector. We also find that pairing-aware metrics explain retrieval better than distributional distances alone. LAYERSCOPE therefore offers a framework for comparing representations across models and layers, enabling a more systematic evaluation in video and multimodal settings.
Figures & tables
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Model Family | Model | Checkpoint | Modalities | Architecture |
| Predictive video | V-JEPA 2 | facebook/vjepa2-vitl-fpc64-256 | Video | ViT-L |
| Audio-video contrastive | PE-AV Small | facebook/pe-av-small | Text/video | Video + text encoders |
| Audio-video contrastive | PE-AV Large | facebook/pe-av-large | Text/video | Video + text encoders |
| Video-text contrastive | X-CLIP | microsoft/xclip-base-patch32 | Text/video | Base dual encoder |
| Autoregressive multimodal | Gemma 4 | google/gemma-4-E4B-it | Multimodal | E4B-IT decoder |
| Autoregressive multimodal | Qwen3-VL | Qwen/Qwen3-VL-Embedding-2B | Text/video | Embedding-2B decoder |
| Analysis | Dataset | Evaluated representations | Evaluation | Score |
| Classification | Breakfast | Every video layer; model-default output | Eight-shot linear probe | Accuracy |
| Classification | UCF101 | Every video layer; model-default output | Eight-shot linear probe | Accuracy |
| Classification | HMDB51 | Every video layer; model-default output | Eight-shot linear probe | Accuracy |
| Clustering | UCF101 | Every video layer; model-default output | MiniBatchKMeans | V-measure |
| Text-to-video retrieval | TUNA-Bench | Every text–video layer pair; model-default pair | Cosine ranking | Mean nDCG@10 |
| Text-to-video retrieval | VATEX | Every text–video layer pair; model-default pair | Cosine ranking | Mean nDCG@10 |
| Task | Model | Matrix | Oracle | Oracle score | Matched depth | Model-default | ||
| TUNA | PE-AV Small | 0.9681 | 0.9681 | +0.0000 | 0.9671 | +0.0010 | ||
| VATEX | PE-AV Small | 0.7914 | 0.7914 | +0.0000 | 0.7899 | +0.0016 | ||
| TUNA | PE-AV Large | 0.9695 | 0.9695 | +0.0000 | 0.9695 | +0.0000 | ||
| VATEX | PE-AV Large | 0.8172 | 0.8172 | +0.0000 | 0.8172 | +0.0000 | ||
| TUNA | X-CLIP | 0.4014 | 0.3329 | +0.0685 | 0.3328 | +0.0686 | ||
| VATEX | X-CLIP | 0.4148 | 0.3950 | +0.0198 | 0.3950 | +0.0198 |
| Model | Best | Best accuracy | Model-default | Difference |
| PE-AV Small | 0.7449 | 0.7366 | +0.0082 | |
| PE-AV Large | 0.8421 | 0.8390 | +0.0031 | |
| X-CLIP | 0.7598 | 0.7603 | ||
| Qwen3-VL | 0.8724 | 0.8719 | +0.0005 | |
| LCO | 0.8323 | 0.8302 | +0.0021 |
| Model | Task | Layers | ||||||
| V-JEPA 2 | Breakfast cls. | 24 | 22 | 0.3110 | +0.0215 | 0.3124 | +0.0201 | |
| V-JEPA 2 | UCF101 cls. | 24 | 23 | 0.8072 | +0.0058 | 0.8062 | +0.0068 | |
| V-JEPA 2 | UCF101 clust. | 24 | 23 | 0.7702 | +0.0006 | 0.7830 | ||
| PE-AV Small | Breakfast cls. | 4 | 4 | 0.5185 | +0.0000 | 0.4529 | +0.0656 | |
| PE-AV Small | UCF101 cls. | 4 | 1 | 0.9248 | +0.0090 | 0.8759 | +0.0579 | |
| PE-AV Small | UCF101 clust. | 4 | 1 | 0.7589 | +0.0694 | 0.8508 |
| Model | ||||||
| V-JEPA 2 | 23 | 0.4690 | 0.4682 | 0.4578 | +0.0009 | +0.0113 |
| PE-AV Small | 1 | 0.6284 | 0.5987 | 0.4247 | +0.0297 | +0.2036 |
| Gemma 4 | 14 | 0.4724 | 0.3904 | 0.0891 | +0.0820 | +0.3833 |
| PE-AV Large | 2 | 0.6462 | 0.6237 | 0.5136 | +0.0225 | +0.1326 |
| X-CLIP | 12 | 0.4807 | 0.4807 | 0.6242 | +0.0000 | |
| Qwen3-VL | 28 | 0.6224 | 0.6224 | 0.6602 | +0.0000 |
| Layer choice | Mean | Median | Maximum | Zero | W/T/L vs. last | |
| Maximum RankMe | 21 | 0.0143 | 0.0081 | 0.0558 | 3 | 8/9/4 |
| Last layer | 21 | 0.0249 | 0.0090 | 0.1694 | 4 | – |
| Maximum standardized covariance effective rank | 21 | 0.0399 | 0.0147 | 0.2172 | 3 | 10/3/8 |
| Minimum GRIDS-LID ( ) | 21 | 0.0447 | 0.0063 | 0.2380 | 4 | 7/4/10 |
| Maximum GRIDS-LID ( ) | 21 | 0.0997 | 0.0374 | 0.4734 | 0 | 3/6/12 |