cs.HCOct 4, 2026

Human-Like Attention? A Psychophysical Comparison of Visual Search in Humans and MLLMs

Authors: Renchi Zhang, Joost C. F. de Winter, Dimitra Dodou, Harleigh C. Seyffert, Yke Bauke Eisma

Organizations: Faculty of Mechanical Engineering, Delft University of Technology, Delft, The Netherlands

Abstract

Visual search is a fundamental cognitive ability. This study investigates whether Multimodal Large Language Models (MLLMs) exhibit human-like difficulty signatures in visual search tasks. We compared search performance of humans (n = 1,250) and MLLMs using identical 2D and 3D stimuli across different set sizes. Both groups showed efficient performance in feature searches, most clearly when the target had a unique color, but performance degradation in conjunction searches as set sizes increased. Additionally, we found strong correlations between human and MLLM error rates (ρ=0.82ρ= 0.82), which suggests that MLLMs are sensitive to similar objective complexities, such as stimulus heterogeneity. However, differences were found as well: whereas humans invested extra search time to respond accurately on target-absent trials, MLLMs exhibited extreme present/absent response biases in complex searches. We conclude that MLLMs replicate high-level human performance signatures, yet their underlying computations differ significantly.

Explore similar work

CardsList
  1. BabyVision: Visual Reasoning Beyond Language

    Jan 10, 2026Liang Chen, Weichu Xie, Yiyan Liang +27Recent Vision-Language ModelsVisual Reasoning

  2. VisLens: Single-Pass Interpretable Visual Search for Multimodal LLMs

    Aug 31, 2026Jingyi He, Sanghwan Kim, Zeynep AkataMultiple Vision TasksRecent Vision-Language Models