cs.CVFeb 12, 2026

Modeling The Object Representations Underlying Human Physical Reasoning

Authors: Andrey Gizdov, Andrea Procopio, Lorenzo Caputi, Georgi I. Ivanov, Yichen Li, Daniel Harari, Tomer Ullman

Organizations: Harvard University · Weizmann Institute of Science · Bocconi University

Abstract

Humans appear to represent objects when reasoning about physics with coarse, volumetric "bodies" that smooth concavities, trading fine visual detail for efficient physical predictions. Yet, the structure of these representations remains largely unknown. Segmentation models, in contrast, are trained for pixel-accurate masks that may misalign with such bodies. We ask whether and when these models nonetheless acquire human-like object representations. Using a time-to-collision (TTC) and change detection (CD) behavioral task with data from 178 and 50 human participants, respectively, we introduce a pipeline and an alignment metric to compare the visual representations of segmentation models to those of humans. We do this systematically on multiple architectures (DINOv2, SegFormer, DeepLabV3+, and UPerNet), varying their size and training time. We find that briefly trained models segment objects too coarsely, aligning poorly with humans, while fully trained models segment objects too finely. For each model, there is an intermediate training regime that best matches the coarse bodies observed in human behaviour, and larger models tend to reach it earlier. We show these bodies emerge under resource constraints in general-purpose vision models, providing computational support to resource-rational accounts of human cognition. This work provides a foundational framework for testing alignment between vision models and humans and shows there is a growing gap between the state-of-the-art in artificial intelligence and human cognition, driven by scaling model size and training.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Human-like Object Grouping in Self-supervised Vision Transformers

    Mar 14, 2026Hossein Adeli, Seoyoung Ahn, Andrew Luo +3Self-Supervised Vision TransformersSelf-Supervised Learning

  2. More Accurate, Less Human: Gestalt Grouping in Vision Models

    Aug 10, 2026Sudhanva Manjunath Athreya, Sai Phani Kumar MalladiVision Foundation Models

  3. Texture Representations in Deep Vision Models: Comparing CNNs, Vision Transformers, and Human Perception

    Jul 9, 2026Ludovica de Paolis, Marco Baroni, Alessandro Laio +1Visual PerceptionConvolutional Neural Networks