Geometric Similarity in VLM Low-Level Vision Representations
Organizations: Florida State University · UNC Chapel Hill
Abstract
Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusion transformers (DiTs). Yet, adapting them efficiently for all-in-one low-level image restoration remains a challenge. Crucially, the field lacks an understanding of how VLMs organize hidden-layer representations and whether these structurally distinct paradigms share a common geometric organization for pixel-level perception. Such shared organization is a prerequisite for building highly transferable, unified restoration VLMs and adapters. In this paper, we systematically investigate representational similarity across 24 low-level tasks spanning 5 categories. We propose GeoSim, a unified four-level framework that analyzes task-conditioned representations from global similarity, local geometry, sparse feature decomposition, and topological verification perspectives. Our formulation applies to the analysis of hidden states in AR models and feature maps in DiTs across same- and cross-task/model settings. Our results reveal the organizing principles of low-level visual representations while exposing their limits in cross-task and cross-model agreement. Ultimately, GeoSim provides an interpretability lens for probing latent transferability in low-level vision and diagnosing model limitations in task- or model-specific scenarios.
Figures & tables
| Symbol | Meaning |
| Input image | |
| Samples per task in each sampling repeat ( ); number of sampling repeats ( ) | |
| Image indices ( ); sampling-repeat index ( ) | |
| Task indices ( for pairs, for single task) | |
| Layer index ( : stable layer); feature dimension ( ); retained principal-component dimension ( ) | |
| Representation of image for task at layer |
| Family | Tasks |
| Restoration | Artifact Removal, Defocus Deblurring, Motion Deblurring, Dehazing, Demoiréing, Denoising, Deraining, Desnowing, Underwater Restoration |
| Removal | Lens Flare Removal, Raindrop Removal, Reflection Removal, Shadow Removal |
| Generation/Enhancement | Colorization, Harmonization, Inpainting, Light Enhancement, Style Transfer, Edge Detection |
| Reconstruction | Super-Resolution, HDR Reconstruction |
| Photometric Correction | Relighting, White Balance Correction, Contrast Enhancement |
| Scale | Models |
| Base ( 7–8B) | Emu3-Chat (8B) ( Wang et al., 2024 ; Wang et al., 2026 ) , Anole-7B ( Chern et al., 2024 ) , Janus-Pro-7B ( Chen et al., 2025 ) , InternVL3.5-8B ( Wang et al., 2025 ) , Qwen3-VL-8B ( Bai et al., 2025 ) |
| Large ( 13B) | Qwen-Image-Edit ( 20B) ( Wu et al., 2025 ) , Emu3.5-Image ( 34B) ( Cui et al., 2025 ) |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Condition | Setup | Level(s) | Metrics |
| Same Task, Same Model | One task and model, with images varied | L1 (primary) L2 (supplementary) | Intra-task cosine (Eq. ( 2 )), Frobenius distance (Eq. ( 24 )), LID (Eqs. ( 12 )–( 13 )) |
| Same Task, Different Models | One task with models varied; identically ordered images and independent SAEs | L1–L3 (primary) L4 (supplementary) | dCor (Eq. ( 3 )), NNGS (Eq. ( 4 )), PCA+PAR (Eq. ( 5 )), TSI (Eq. ( 7 )) and feature–task MI (Eq. ( 8 )), RBF-MMD (Eq. ( 9 )), -Wasserstein distance (Eq. ( 21 )), bottleneck distance (Eq. ( 22 )) |
| Different Tasks, Same Model | One model with tasks varied | L1 (primary) | Inter-task cosine (Eq. ( 1 )), MDS (visualization) |
| Different Tasks, Different Models | Tasks and models varied; no image-ID or SAE-coordinate alignment | L3 (primary) L4 (supplementary) | Task-conditioned activation-rate MMD extension of Eq. ( 9 ), averaged reciprocally and then ranked within each model pair; -Wasserstein and bottleneck distances (Eqs. ( 21 ) and ( 22 )) |