cs.CVOct 1, 2026

Geometric Similarity in VLM Low-Level Vision Representations

Authors: Shao-Jun Xia, Huixin Zhang, Zhen Lei, Anlan Sun, Yuner Zhang, Xiaoyang Chen

Organizations: Florida State University · UNC Chapel Hill

Abstract

Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusion transformers (DiTs). Yet, adapting them efficiently for all-in-one low-level image restoration remains a challenge. Crucially, the field lacks an understanding of how VLMs organize hidden-layer representations and whether these structurally distinct paradigms share a common geometric organization for pixel-level perception. Such shared organization is a prerequisite for building highly transferable, unified restoration VLMs and adapters. In this paper, we systematically investigate representational similarity across 24 low-level tasks spanning 5 categories. We propose GeoSim, a unified four-level framework that analyzes task-conditioned representations from global similarity, local geometry, sparse feature decomposition, and topological verification perspectives. Our formulation applies to the analysis of hidden states in AR models and feature maps in DiTs across same- and cross-task/model settings. Our results reveal the organizing principles of low-level visual representations while exposing their limits in cross-task and cross-model agreement. Ultimately, GeoSim provides an interpretability lens for probing latent transferability in low-level vision and diagnosing model limitations in task- or model-specific scenarios.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. GeoWorld-VLM: Geometry from World Models for Vision-Language Models

    May 15, 2026Renjie Gu, Kaichen Zhou, Yan Luo +1Spatial SupervisionVideo World Models

  2. Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision

    Date pendingYuandong Pu, Le Zhuo, Kaiwen Zhu +7Computer VisionRestoration

  3. Same Answer, Different Representations: Hidden instability in VLMs

    Feb 6, 2026Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena +6InstabilityAnswer