Human face perception reflects inverse-generative and naturalistic discriminative objectives
Authors: Wenxuan Guo, Heiko H. Schütt, Kamila Maria Jozwik, Katherine R. Storrs, Nikolaus Kriegeskorte, Tal Golan
Organizations: Department of Psychology, Columbia University, New York, NY, USA. · Department of Behavioural and Cognitive Sciences, Université du Luxembourg, Esch-sur-Alzette, Luxembourg. · MRC Cognition and Brain Sciences Unit, University of Cambridge, Cambridge, England. · School of Psychology, University of Auckland, Auckland, New Zealand. · Department of Neuroscience, Columbia University, New York, NY, USA. · Department of Industrial Engineering and Management, Ben-Gurion University of the Negev, Be’er Sheva, Israel. · School of Brain Sciences and Cognition, Ben-Gurion University of the Negev, Be’er Sheva, Israel.
The perceptual representations supporting our ability to recognize faces remain a computational mystery. Deep neural networks offer mechanistic hypotheses for human face perception, but theoretically distinct models often make indistinguishable representational predictions for randomly sampled faces. To expose diagnostic differences among these hypotheses, we compared six neural network models sharing an architecture but trained on distinct tasks, using face pairs optimized to elicit contrasting model predictions ("controversial" pairs) alongside randomly sampled pairs. We tested model predictions against face-dissimilarity judgments from 864 human participants across stimulus sets differing in realism and pose variation. Models prioritizing high-level, invariant structures (trained via inverse rendering, face identification, or object classification) most robustly matched human judgments. Furthermore, models trained on natural images typically outperformed synthetic-trained counterparts. Together, these findings suggest that human face perception is shaped by mechanisms that infer latent causes of facial appearance, discount nuisance variation, and are tuned by natural image statistics.