Understanding how learned representations respond to finite input changes is important for characterizing their sensitivity, invariances, and robustness. Yet existing geometric analyses are predominantly local and describe only infinitesimal perturbations. We introduce a scale-resolved statistic that compares an encoder's measured feature displacement with its local linear prediction as the perturbation magnitude increases. Across diverse image encoders, we discover a characteristic plateau-rise-peak-decay profile, which we call the bump. The bump is absent at initialization, emerges early during standard training, and does not form under randomized labels or random-noise inputs. Its shape also varies with the training distribution and robustness objective. These results establish departures from local geometry as a signature of how encoder representations are shaped by learning.
Figures & tables
Figure 1: Scale-resolved response profiles. The horizontal axis is the perturbation magnitude η , and the vertical axis is rη , the ratio of the measured feature displacement to its local linear prediction. (a) Individual profiles of single images are thin and their median is bold (ConvNeXt-Tiny encoder). (b) Across diverse encoder architectures, trained models (solid) exhibit the characteristic plateau–rise–peak–decay bump , whereas randomly initialized models (dashed) do not (medians reported). (c) During training, the bump forms early and remains largely stable.
Figure 2: Response profiles under controlled training conditions. (a) Adversarial training moves the peak 10 – 30× outward, primarily through an approximately 2000× reduction in the local prediction. (b) ResNet-18 trained on clean or blurred ( σ∈{1,2} ) CIFAR-10 images evaluated on clean or blurred test images. The bump mostly grows with evaluation blur, and shrinks with training blur. (c) Profiles for ResNet-18 trained on CIFAR-10 with randomized labels and random data show no peak on any evaluation set.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Per-image response profiles for the trained ImageNet encoders (rows) on the five ImageNet-resolution evaluation sets (columns). Each line is one image. The plateau–rise–peak–decay shape persists across encoders and evaluation sets; height and location vary.
Figure 4: Per-image response profiles for the randomly initialized controls. No interior peak forms on any evaluation set: RN-50 and ViT-S/16 rise monotonically to the end of the ladder, and ConvNeXt-Tiny stays near rη=1 before decaying
Figure 5: Per-image response profiles for the adversarially trained ResNet-50 family. With increasing training budget ε , the peak moves to larger η and grows; the baseline ε=0 baseline retains a peak at small scale.
Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object-background or texture correlations. We identify a different source of shortcut learning: invisible metadata traces embedded at the pixel level, for metadata such as image processing and photo acquisition. We hypothesize that large-scale semantic supervision, whether through categorical labels (ImageNet) or billion-scale captions (LAION), naturally induces metadata-semantics correlations during pretraining, leading models to convert low-level signals into predictive features. By introducing controlled metadata-semantics correlations, we show that stronger ones produce systematically higher sensitivity to metadata traces and larger performance degradation under metadata distribution shifts. We further explore mitigation strategies applied during and after pretraining that reduce sensitivity not only to targeted metadata but also to unseen ones, without sacrificing performance on downstream tasks. Metadata sensitivity also has a positive side: it partly explains the strong generated-image detection ability of some encoders, while its mitigation can improve out-of-distribution generalization. Code: https://github.com/ryan-caesar-ramos/visual-encoder-traces
Vladan Stojnić, Ryan Ramos, Giorgos Kordopatis-Zilos +2
VRG, FEE, Czech Technical University in Prague · The University of Osaka
Representational similarity methods compare the geometries of neural representations, but they do not measure how consistently the geometry of a single representation is recovered from subsets of its feature coordinates. We call this property geometric stability and introduce Shesha, which estimates it by correlating representational dissimilarity matrices from complementary random feature subsets. Shesha is not invariant to orthogonal rotations: representations with identical Gram matrices, and therefore identical linear CKA, can have different geometric stability. Controlled transformations further separate the quantities. Across 2,463 encoder configurations spanning seven domains, similarity and stability are positively associated across non-PCA transformations (ρ=+0.75) but negatively associated under PCA-coordinate compression (ρ=−0.47). We further evaluate 170 pretrained vision models across six datasets. DINOv2 combines strong transfer performance with bottom-quartile stability on five of six datasets, showing that transferability and feature-split stability need not coincide. Across random feature subsets, the marginal relationship between Shesha and linear-probe variability is dataset-dependent; after controlling for task alignment with LogME, higher Shesha is associated with lower variability on five of six datasets. These results identify geometric stability as a basis-dependent property that complements representational similarity and task alignment.
Intermediate feature representations represent the backbone for the expressivity and adaptability of deep neural networks. However, their geometric structure remains poorly understood. In this submission, we provide indirect insights into this matter by applying a broad selection of manipulations in input space, ranging from geometric and photometric transformations to local masking and semantic manipulations using generative image editing models, and assess the feasibility of learning a mapping in the feature space, mapping from the original to the manipulated feature map. To this end, we devise different types of mappings, from linear to non-linear and local to global mappings and assess both the reconstruction quality of the mapping as well as the semantic content of the mapped representations. We demonstrate the feasibility of learning such mappings for all considered transformations. While global (transformer) models that operate on the full feature map often achieve best results, we show that the same can be achieved with a shared linear model operating on a single feature vector typically with very little degradation in reconstruction quality, even for highly non-trivial semantic manipulations. We analyze the corresponding mappings across different feature layers and characterize them according to dominance of weight vs. bias and the effective rank of the linear transformations. These results provide hints for the hypothesis that the feature space is to a first degree of approximation organized in linear structures. From a broader perspective, the study demonstrates that generative image editing models might open the door to a deeper understanding of the feature space through input manipulation.
Elias B. Krey, Nils Neukirch, Nils Strodthoff
Division AI4Health Carl von Ossietzky Universität Oldenburg Oldenburg, Germany