Understanding how learned representations respond to finite input changes is important for characterizing their sensitivity, invariances, and robustness. Yet existing geometric analyses are predominantly local and describe only infinitesimal perturbations. We introduce a scale-resolved statistic that compares an encoder's measured feature displacement with its local linear prediction as the perturbation magnitude increases. Across diverse image encoders, we discover a characteristic plateau-rise-peak-decay profile, which we call the bump. The bump is absent at initialization, emerges early during standard training, and does not form under randomized labels or random-noise inputs. Its shape also varies with the training distribution and robustness objective. These results establish departures from local geometry as a signature of how encoder representations are shaped by learning.
Figures & tables
Figure 1: Scale-resolved response profiles. The horizontal axis is the perturbation magnitude η , and the vertical axis is rη , the ratio of the measured feature displacement to its local linear prediction. (a) Individual profiles of single images are thin and their median is bold (ConvNeXt-Tiny encoder). (b) Across diverse encoder architectures, trained models (solid) exhibit the characteristic plateau–rise–peak–decay bump , whereas randomly initialized models (dashed) do not (medians reported). (c) During training, the bump forms early and remains largely stable.
Figure 2: Response profiles under controlled training conditions. (a) Adversarial training moves the peak 10 – 30× outward, primarily through an approximately 2000× reduction in the local prediction. (b) ResNet-18 trained on clean or blurred ( σ∈{1,2} ) CIFAR-10 images evaluated on clean or blurred test images. The bump mostly grows with evaluation blur, and shrinks with training blur. (c) Profiles for ResNet-18 trained on CIFAR-10 with randomized labels and random data show no peak on any evaluation set.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Per-image response profiles for the trained ImageNet encoders (rows) on the five ImageNet-resolution evaluation sets (columns). Each line is one image. The plateau–rise–peak–decay shape persists across encoders and evaluation sets; height and location vary.
Figure 4: Per-image response profiles for the randomly initialized controls. No interior peak forms on any evaluation set: RN-50 and ViT-S/16 rise monotonically to the end of the ladder, and ConvNeXt-Tiny stays near rη=1 before decaying
Figure 5: Per-image response profiles for the adversarially trained ResNet-50 family. With increasing training budget ε , the peak moves to larger η and grows; the baseline ε=0 baseline retains a peak at small scale.