Humans appear to represent objects when reasoning about physics with coarse, volumetric "bodies" that smooth concavities, trading fine visual detail for efficient physical predictions. Yet, the structure of these representations remains largely unknown. Segmentation models, in contrast, are trained for pixel-accurate masks that may misalign with such bodies. We ask whether and when these models nonetheless acquire human-like object representations. Using a time-to-collision (TTC) and change detection (CD) behavioral task with data from 178 and 50 human participants, respectively, we introduce a pipeline and an alignment metric to compare the visual representations of segmentation models to those of humans. We do this systematically on multiple architectures (DINOv2, SegFormer, DeepLabV3+, and UPerNet), varying their size and training time. We find that briefly trained models segment objects too coarsely, aligning poorly with humans, while fully trained models segment objects too finely. For each model, there is an intermediate training regime that best matches the coarse bodies observed in human behaviour, and larger models tend to reach it earlier. We show these bodies emerge under resource constraints in general-purpose vision models, providing computational support to resource-rational accounts of human cognition. This work provides a foundational framework for testing alignment between vision models and humans and shows there is a growing gap between the state-of-the-art in artificial intelligence and human cognition, driven by scaling model size and training.
Figures & tables
Figure 1: Research overview. (A) Different object representations are useful for different goals. “Body” representation is useful for physical reasoning. “Shape” representation is useful for recognition. (B) Prior work ( Li et al., 2023 ) has shown evidence that humans do have such coarse body representations for physical reasoning (e.g. predicting collision times), specifically by “filling in” concave regions. (C) Our research asks whether vision segmentation models have such similar representations and under what conditions they develop them.
Figure 2
Figure 3: Alignment with the human TTC effect peaks at intermediate training. Eˉ (ms) across training, averaged over seeds; bands show ± SEM, and darker lines are larger models. Eˉ stays above zero because the absolute value is taken per seed before averaging (Appendix D.3 ).
Figure 4: Larger models reach low alignment error earlier. Eˉ (ms) over training steps (horizontal) and model size (vertical) for each family. Low Eˉ forms a band that occurs earlier in training for larger models.
Figure 5: Concavities are resolved after convex parts. Masks (top) and foreground softmax probability maps (bottom) for SegFormer-B3 and DeepLabV3+-R50 at an early, an intermediate, and a late training step. Convex parts sharpen first; concave notches stay partly filled at intermediate steps and are resolved late.
Figure 6: Models are less sensitive to concave than to convex changes, like humans. (A) RAC across training; darker lines are larger models, and the dotted line marks RAC =1 (the added area is fully captured). Concave RAC is below convex RAC for most of training. (B) Detection rates with one threshold ( τ=1.1% ) fitted on SegFormer-B5 and applied to all configurations shown.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Change-detection stimuli. The same area is added either inside a concavity (orange) or on a convex part (green).
Family
Backbones
Params (M)
Initialization
SegFormer ( Xie et al., 2021 )
B0–B5
3.7–84.6
ADE20K ( Zhou et al., 2017 )
DeepLabV3+ ( Chen et al., 2018 )
ResNet-18/50/101
12.3–60.4
Cityscapes ( Cordts et al., 2016 )
UPerNet ( Xiao et al., 2018 )
ConvNeXt S/B/L/XL
82.0–391.0
ADE20K
DINOv2 † ( Oquab et al., 2024 )
S/B/L
27.2–310.6
Self-supervised
Appendix
Table 1: Model families. Parameter ranges span the smallest to the largest backbone in each family. † With a light convolutional decoder.
Setting
Value
Optimizer
AdamW (weight decay 0.01)
Learning rate
5×10−5
Scheduler
Cosine annealing with linear warmup (5% of total steps)
Loss
Cross-entropy + Dice ( λ=0.5 )
Effective batch size
8 (gradient accumulation)
Training length
15 epochs (945 steps)
Appendix
Table 2: Fine-tuning hyperparameters, shared by all 16 configurations.
Figure 8: Masks across model size, at a fixed training step (step 144 for SegFormer B0/B2/B4; step 118 for DeepLabV3+ R18/R50/R101). Top: mask overlays. Bottom: foreground probability maps, obtained applying a softmax to the logits.
Model
Params (M)
Best step
Eˉ (ms)
Δ (ms)
Pairwise
Test
SegFormer B0
3.7
202
2.4±0.5
31.3±1.8
45/48
SegFormer B1
13.7
164
5.3±1.4
32.4±3.3
43/48
SegFormer B2
27.4
102
5.0±3.4
28.6±4.3
44/48
SegFormer B3
47.2
98
4.9±2.7
27.0±3.1
46/48
SegFormer B4
64.1
92
3.4±1.3
31.0±2.3
46/48
SegFormer B5
84.6
102
3.0±1.8
32.1±2.5
46/48
Appendix
Table 3: TTC statistics at minimum-error checkpoints (in-sample). The best configuration is in bold. The human row gives the mean effect and its 95% CI.
Aligned < first
Aligned < last
Size slope
Family
n
t
p
t
p
Δ in human CI
β^
p
SegFormer
6
−2.68
.022
−27.4
<.0001
6/6
−0.037
.020
UPerNet
4
−4.15
.013
−8.03
.002
4/4
−0.090
.216
DINOv2
3
−2.54
.063
−31.6
<.001
3/3
−0.197
.073
DeepLabV3+
3
−5.25
.017
−12.6
.003
3/3
−0.012
.308
All
16
−5.17
5.7×10−5
−19.6
2.1×10−12
16/16
–
–
Appendix
Table 4: Per-family tests of the training trajectory. The aligned checkpoint is selected on one half of the participants, and Eˉ is compared on the other half with the first and last checkpoints (one-sided paired t -tests across the n backbones of each family, df=n−1 ). Δ in human CI: configurations whose seed-averaged concavity effect lies inside the human 95% CI for a window of checkpoints. β^ : OLS slope of model-size rank on the minimum- Eˉ training step, in size ranks per step; negative values mean larger models reach the minimum earlier (two-sided p , df=n−2 ; Appendix D.4 ).
Figure 9: Signed TTC concavity effect across training. Positive values mean earlier predicted collision for concave than for convex configurations. Lines: mean across seeds; bands: ± SEM; dashed line: human effect.
Operator
Fitted parameter
Eˉ (ms)
Morphological closing
r=36
1.6
α -shape
α=24
3.2
Dilation
r=6
4.8
Convex hull
none
412.8
Models, minimum-error checkpoints (Table 3 )
0.8–14.5
Appendix
Table 5: Geometric operators applied to ground-truth masks, with parameters fitted to the human data and scored with the same Eˉ as the models.
Figure 10: Detection rate as a function of the threshold τ on δ , at the final training step, averaged over backbones and seeds. Solid lines: model detection rates for concave (yellow) and convex (green) changes. Dashed lines: human detection rates.
Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent object segmentation properties. However, their alignment with human object perception remains poorly understood. Here, we introduce a behavioral benchmark in which participants make same/different object judgments for dot pairs on naturalistic scenes, scaling up a classical psychophysics paradigm to over 1000 trials. We test a diverse set of vision models using a simple readout from their representations to predict subjects' reaction times. We observe a steady improvement across model generations, with both architecture and training objective contributing to alignment, and transformer-based models trained with the DINO self-supervised objective showing the strongest performance. To investigate the source of this improvement, we propose a novel metric to quantify the object-centric component of representations by measuring patch similarity within and between objects. Across models, stronger object-centric structure predicts human segmentation behavior more accurately. We further show that matching the Gram matrix of supervised transformer models, capturing similarity structure across image patches, with that of a self-supervised model through distillation improves their alignment with human behavior, converging with the prior finding that Gram anchoring improves DINOv3's feature quality. Together, these results demonstrate that self-supervised vision models capture object structure in a behaviorally human-like manner, and that Gram matrix structure plays a role in driving perceptual alignment.
Hossein Adeli, Seoyoung Ahn, Andrew Luo +3
Zuckerman Mind Brain Behavior Institute, Columbia University, New York · Department of Social Science and AI, Hankuk University of Foreign Studies, Seoul · University of Hong Kong, Hong Kong +2
Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into recognizable objects. These are the Gestalt operations that visualization design builds on. Whether vision models organize visual content this way has not been systematically tested. We introduce a behavioral battery that scores models against human data from prior perception studies on four grouping tasks: mark-color odd-one-out, color-series counting, silhouette recognition, and object odd-one-out. We apply it to 45 models across five training families: supervised, self-supervised, and contrastive vision-language encoders, open-weight VLMs, and closed foundation models. The battery reveals that agreement with human responses captures aspects of perceptual organization that conventional performance metrics fail to distinguish, with several closed models exhibiting substantially lower alignment than their benchmark accuracy would suggest. Scoring against published perception data therefore gives visualization research a reusable yardstick, requiring no new user study, for auditing whether the models now entering visualization pipelines organize what they see the way their human audience does.
Sudhanva Manjunath Athreya, Sai Phani Kumar Malladi
In computational vision science, Convolutional Neural Networks (CNNs) have emerged as a popular model of biological vision because of the alignment they can exhibit with neural and behavioral data in humans and animals. However, it remains unclear to what extent this alignment persists for visual tasks that extend beyond the canonical object recognition paradigm based on well defined semantic content. In this study, we diverge from the common object-centric view by focusing on another aspect of vision: texture perception. We consider textures of different complexity generated with three different algorithms from the same source images. Using a rank-based statistic, we quantify the information encoded in the internal representations of a CNN and three Vision Transformers (ViTs), and we compare the similarity of these representations to those inferred from human psychophysics data. We find that the representation of textures is aligned in different ViTs, but not between the ViTs and the CNN; that ViTs form similar representations for textures of different complexity; that human performance in recognizing textures can be better predicted from ViTs representations rather than CNN representations. Taken together, these results suggest that ViTs may capture more faithfully than CNNs how texture patterns are visually processed by humans, and that the representations of texture stimuli in computational models may be driven by the network architecture.
Ludovica de Paolis, Marco Baroni, Alessandro Laio +1
Department of Neuroscience, International School for Advanced Studies (SISSA), Trieste, Italy · Department of Language and Translation Sciences, Pompeu Fabra University, Barcelona, Spain · ICREA, Barcelona, Spain +1