Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusion transformers (DiTs). Yet, adapting them efficiently for all-in-one low-level image restoration remains a challenge. Crucially, the field lacks an understanding of how VLMs organize hidden-layer representations and whether these structurally distinct paradigms share a common geometric organization for pixel-level perception. Such shared organization is a prerequisite for building highly transferable, unified restoration VLMs and adapters. In this paper, we systematically investigate representational similarity across 24 low-level tasks spanning 5 categories. We propose GeoSim, a unified four-level framework that analyzes task-conditioned representations from global similarity, local geometry, sparse feature decomposition, and topological verification perspectives. Our formulation applies to the analysis of hidden states in AR models and feature maps in DiTs across same- and cross-task/model settings. Our results reveal the organizing principles of low-level visual representations while exposing their limits in cross-task and cross-model agreement. Ultimately, GeoSim provides an interpretability lens for probing latent transferability in low-level vision and diagnosing model limitations in task- or model-specific scenarios.
Figures & tables
Figure 1: Overview of VLM representations for low-level tasks and the GeoSim framework.
Symbol
Meaning
x
Input image
n,b
Samples per task in each sampling repeat ( 100 ); number of sampling repeats ( 10 )
i,j,s
Image indices ( i,j ); sampling-repeat index ( s )
p,q,t∈T
Task indices ( p,q for pairs, t for single task)
ℓ,d,dPCA
Layer index ( ℓ∗ : stable layer); feature dimension ( d ); retained principal-component dimension ( dPCA )
Relighting, White Balance Correction, Contrast Enhancement
Table 2: Low-level vision task suite organized by task family.
Scale
Models
Base ( ∼ 7–8B)
Emu3-Chat (8B) ( Wang et al., 2024 ; Wang et al., 2026 ) , Anole-7B ( Chern et al., 2024 ) , Janus-Pro-7B ( Chen et al., 2025 ) , InternVL3.5-8B ( Wang et al., 2025 ) , Qwen3-VL-8B ( Bai et al., 2025 )
Large ( > 13B)
Qwen-Image-Edit ( ∼ 20B) ( Wu et al., 2025 ) , Emu3.5-Image ( ∼ 34B) ( Cui et al., 2025 )
Table 3: VLMs used in our analyses.
Figure 2: Condition I: Consistency of the instruction-induced displacement across network depth. Each panel is one model; each of the 24 curves is one task, colored by perceptual family.
Figure 3: Condition II: Task-conditioned cross-model and within-model reference MMD (left), per-model TSI composition (center), and feature–task MI (right).
Figure 7
Figure 6: Condition IV: Cross-task, cross-model activation-rate distribution similarity from the Level-3 sweep. Results are aggregated using MMD across seven models and 24 tasks.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Feature extraction and stable-layer selection. Illustrated using a typical AR model (Emu3). Input images with task and neutral prompts produce the cached states whose difference defines hpℓ(x) .
Figure 8: SAE training and its use in Conditions II and IV. Per-model SAEs are fit on task-balanced training data and reused to compare held-out activation-rate distributions. Neither condition assumes aligned SAE feature coordinates.
Tasks and models varied; no image-ID or SAE-coordinate alignment
L3 (primary) L4 (supplementary)
Task-conditioned activation-rate MMD extension of Eq. ( 9 ), averaged reciprocally and then ranked within each model pair; H0/H1r -Wasserstein and bottleneck distances (Eqs. ( 21 ) and ( 22 ))
Appendix
Table 4: Task/model conditions, and corresponding analysis levels and metrics.
Figure 9: Condition I: Same Task, Same Model. LID per task, estimated with the Levina–Bickel estimator at k=10 at each model’s analysis layer ℓ∗ . Tasks are ordered by their cross-model mean, shown as the black line.
Figure 10: Condition II: Same Task, Different Models. Local dimensionality against cross-model agreement, one point per task. The abscissa is the seven-model mean LID of Figure 9 . The ordinates are the held-out PAR (left) and the chance-adjusted neighborhood overlap (right), each averaged over the 21 model pairs. The dotted line marks the PAR baseline of 1.0, the residual obtained by predicting the mean.
Figure 11: Condition II: Same Task, Different Models. Level-4 2-Wasserstein and bottleneck distances for H0 and H1 across 16 model-pair and task combinations.
Figure 12: Condition III: Different Tasks, Same Model. Inter-task cosine similarity matrices, with one panel per model. Both axes carry the same 24 tasks in a fixed order grouped by perceptual family, with rules and labels marking the five categories. Each panel is standardized on its own off-diagonal entries, so color encodes how far a task pair departs from the typical pair of that model, red above and blue below, on the common scale of ±2.2 SD given in the color bar. The diagonal is left blank because it carries intra-task rather than inter-task similarity.
Figure 13: Condition IV: Different Tasks, Different Models. Level-3 screening rank and Level-4 H0/H1 topological distances for the 10 reciprocal candidates, shown on raw and scale-normalized views. Dots show 2-Wasserstein means over the ten sampling runs and both task assignments, thin bars span the bootstrap percentile intervals of both assignments, and thick bars show the range between their means.
Modern Vision-Language Models (VLMs) achieve strong semantic recognition, yet remain brittle on elementary spatial relations such as left of, on, behind, and between. One cause of this failure arises before language reasoning begins: the visual pathway may compress or discard critical 3D structural cues during feature extraction, so the language model receives image representations that are already insufficient for reliable spatial judgment. We introduce GeoWorld-VLM, a VLM-side distillation framework that transfers geometric structure from frozen camera-conditioned video world models into VLMs. GeoWorld-VLM fine-tunes only the image encoder and multimodal projector, aligning post-projector image features with intermediate world-model representations while leaving the main backbone frozen. Given images, a prompt, and a sampled camera trajectory, the world-model teacher converts static visual input into a synthetic multi-view spatial signal. Training combines spatial answer supervision, teacher-student feature alignment, and a preservation anchor to the original VLM. Since the language model remains frozen, GeoWorld-VLM preserves the original model's linguistic capabilities while attributing spatial improvements to the enhanced visual pathway. To evaluate the effectiveness and generality of the proposed method, we apply GeoWorld-VLM to two distinct VLM architectures and observe consistent improvements across both backbones. GeoWorld-VLM improves performance by approximately 4 percent on both the What'sUp and VSR benchmarks, suggesting that world-model-guided visual alignment generalizes across model structures and spatial reasoning datasets.
Renjie Gu, Kaichen Zhou, Yan Luo +1
Harvard AI and Robotics Lab · Kempner Institute for the Study of Natural and Artificial Intelligence · Harvard University
We present Lunima-OmniLV (abbreviated as OmniLV), a universal multimodal multi-task framework for low-level vision that addresses over 100 sub-tasks across four major categories: image restoration, image enhancement, weak-semantic dense prediction, and stylization. OmniLV leverages both textual and visual prompts to offer flexible and user-friendly interactions. Built on Diffusion Transformer (DiT)-based generative priors, our framework supports arbitrary resolutions -- achieving optimal performance at 1K resolution -- while preserving fine-grained details and high fidelity. Through extensive experiments, we demonstrate that separately encoding text and visual instructions, combined with co-training using shallow feature control, is essential to mitigate task ambiguity and enhance multi-task generalization. Our findings also reveal that integrating high-level generative tasks into low-level vision models can compromise detail-sensitive restoration. These insights pave the way for more robust and generalizable low-level vision systems.
Yuandong Pu, Le Zhuo, Kaiwen Zhu +7
Shanghai Jiao Tong University · Shanghai AI Laboratory · University of Macau +3
The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions reflect stable multimodal processing. In this work, we argue that this assumption is insufficient. We introduce a representation-aware and frequency-aware evaluation framework that measures internal embedding drift, spectral sensitivity, and structural smoothness (spatial consistency of vision tokens), alongside standard label-based metrics. Applying this framework to modern VLMs across the SEEDBench, MMMU, and POPE datasets reveals three distinct failure modes. First, models frequently preserve predicted answers while undergoing substantial internal representation drift; for perturbations such as text overlays, this drift approaches the magnitude of inter-image variability, indicating that representations move to regions typically occupied by unrelated inputs despite unchanged outputs. Second, robustness does not improve with scale; larger models achieve higher accuracy but exhibit equal or greater sensitivity, consistent with sharper yet more fragile decision boundaries. Third, we find that perturbations affect tasks differently: they harm reasoning when they disrupt how models combine coarse and fine visual cues, but on the hallucination benchmarks, they can reduce false positives by making models generate more conservative answers.
Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena +6
Sapienza University of Rome · University of Edinburgh · University of Oxford +5