Humans perceive far more in a scene than what is explicitly depicted: a single glance captures past causes and future trajectories; a quick peek determines if a vehicle can fit between two parked cars; a few seconds of video reveals who holds authority in a room; and a fleeting clip highlights subtle abstract patterns like unwritten rules or hidden labels. This capacity reflects a form of humanity's sixth sense: an intuitive reasoning mechanism that recovers implicit information beyond raw sensory perception. Crucially, this rapid, zero-shot visual intuition underpins everyday navigation and social interaction, making it a vital capability for Multimodal Large Language Models (MLLMs) deployed alongside people. Existing visual benchmarks, however, target either deliberate expert-level analysis in academic and mathematical domains or low-level perception, leaving the intuitive reasoning that people perform largely untested. To bridge this gap, we introduce Humanity's Sixth Sense (HSS), a benchmark for intuitive visual reasoning. HSS spans diverse image and video inputs, organizes items under a structured taxonomy, and pairs each with human-written prompts probing the implicit temporal, spatial, social, and abstract structure that people infer at a glance. Frontier MLLMs fall short of human performance: participants reach 93.1% accuracy, while the strongest model, GPT-6-astra, reaches only 53.6% even at maximum reasoning effort. Despite excelling in many complex tasks that require advanced perception and knowledge, current models still struggle significantly on these visual tasks that are intuitive for humans. We further explore agentic setup that apply dynamic visual manipulation to HSS, which narrows but does not close the gap. HSS establishes intuitive visual reasoning as a measurable axis and directs attention to a capability that scaling on current benchmarks has so far left behind.
Figures & tables
Figure 2: HSS task examples across the four domains. Each panel shows the input media (a single image or sampled video keyframes), human-written question, reference answer, and evaluation rubric. Temporal & causal dynamics : identifying the physical outcome of a manual track-switching action in video. Physical & spatial logic : judging whether a vehicle can fit between two parked sedans without explicit scale markers. Social understanding : inferring why a person stops running after a bus in a video clip. Abstract & contextual inference : deducing an occluded beaker label from an arithmetic progression of neighboring legible labels.
Vendor
Model (Reasoning Effort)
pass@1
image
video
temporal
physical
social
abstract
think-tok
ans-tok
Human baseline
93.1
94.1
91.9
97.0
94.9
89.1
89.2
—
—
OpenAI
GPT-6-astra (max)
53.6±4.1
55.6
51.1
58.2
53.6
46.1
55.6
2,896
40
GPT-6.1-sol (max)
46.6±3.8
45.5
47.9
55.2
45.8
36.4
47.4
2,298
36
GPT-6-sol (max)
31.2±3.4
31.8
30.5
38.1
28.6
22.4
36.3
3,298
30
GPT-5.6-sol (max)
30.0±3.3
28.5
31.9
38.6
26.3
27.3
28.1
3,630
39
GPT-5.6-terra (max)
22.7±3.2
24.1
20.9
28.4
21.0
19.1
21.9
5,794
45
Table 1: Main results. Pass@1 (%) on HSS. A task counts as correct only when every rubric criterion is met; ± is the half-width of a 95% bootstrap interval over tasks. The columns re-cut the same tasks by media type and by domain, and think-tok / ans-tok are mean thinking and answer tokens per task. The best model value is bold and the second best is underlined , and the blind control re-runs the strongest model with the media removed.
Figure 3: What models get wrong, by subdomain. Share (%) of each subdomain’s failures attributed to each cause, over 8,573 labelled failures from 24 models on 522 tasks.
Overall pass@1
Δ by media
Δ by domain
Cost
Harness
Base model
raw
agent
Δ
image
video
temporal
physical
social
abstract
think-tok
steps
Claude Code
Claude-Opus-5
30.8
51.3
+20.5
+24.3
+15.8
+26.3
+15.1
+18.2
+25.6
43,764
72
Claude-Fable-5.1
41.2
57.7
+16.5
+15.4
+17.9
+16.8
+14.9
+22.1
+13.2
42,237
57
Codex
GPT-6-astra
54.4
59.3
+4.9
+4.5
+5.5
+6.7
−1.4
+9.5
+9.6
2,040
14
GPT-5.6-sol
28.4
35.8
+7.5
+8.0
+6.8
+3.4
+6.5
+6.9
+15.5
6,541
18
Table 2: Agentic harnesses against single-pass inference . Raw is the base model called once, Agent the same model inside the harness; every Δ is agentic minus single-pass on those tasks, re-cut by media type and domain as in Table 1 . think-tok is mean thinking tokens per task inside the harness and steps the median number of agent steps.
Figure 4: What the agentic scaffold fixes. Pooled over the four harness/base pairs of Table 9 , first attempts. The agent cuts failures from 944 to 759 , mostly on causes a closer look can fix: object/role ( 181→147 ), facing ( 117→93 ), size/fit ( 137→105 ), and missed-cue ( 154→110 ) errors. Depth errors (a 2D overlap read as 3D alignment) do not fall ( 102→101 ), and the small reasoning causes grow.
Figure 5: Performance by subdomains under different reasoning effort We increase the reasoning effort level from left to right. Green is a gain, red a loss.
Figure 6: Judge alignment across vendors. (a) Pass@1 scores under alternative judges closely track the primary judge (Claude-Opus-5) along the identity line. (b) Pairwise agreement exceeds 95% at both the criterion ( n=3,780 ) and the item ( n=2,620 ) level for all three judge pairs.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Media statistics in HSS. (a) Hierarchical breakdown of image and video assets across broad visual categories ( People & social , Work, craft & media , Nature & outdoors , Street & transport , and Home & retail ) and fine-grained sub-scenes. (b) Distribution of visual scene types across the four core reasoning domains, showing diverse environment coverage per domain. (c) Histogram of video clip durations (log scale), spanning a few seconds to 28 minutes; summary statistics are in Table 3 .
Tasks
Mean length (words)
Domain / subdomain
All
Image
Video
Criteria
Prompt
Rubric
Answer
Physical & spatial logic
176
114
62
248
37.6
26.9
73.9
Hidden & invisible
58
28
30
83
27.7
21.8
69.7
Spatial “alien viewpoint”
50
40
10
72
42.0
30.8
78.9
Affordance & feasibility
41
34
7
60
45.4
30.0
77.8
Spatial reachability
27
12
15
33
39.0
25.8
68.1
Appendix
Table 3: HSS dataset composition and statistics. The released benchmark: 522 tasks over four domains and eleven subdomains, graded against 723 rubric criteria ( 1.39 per task; 71.5% carry exactly one). Each task is paired with a unique image or video clip, so task counts are also media counts. The video half totals 17.6 hours: median 76 s, interquartile range 20 – 307 s, longest 28 min. Word counts are means over whitespace-delimited tokens of the prompt, the full task rubric, and the reference answer.
Figure 8: How an HSS task is built and reviewed , with the attrition at each stage. A task is authored in three steps, then faces three independent review rounds; each round may pass it on, return it to authoring for revision (after which it re-enters the pipeline from the start), or drop it outright. Counts are net survivors: a task returned for revision and later re-admitted is counted once, at the round it finally passes, so each dropped figure is the net loss at that round. Of 3,466 authored tasks, 522 survive all three rounds, a yield of 15.1% . The first gate is by far the steepest: it turns away 2,302 tasks, two in three of everything written, overwhelmingly because the answer was already depicted in the scene, or for want of unambiguous consensus on the ground truth (principles 1 and 3 of Section 2.3 ).
Domain
Domain Description
Sub-domain
Sub-domain Description
Temporal & Causal Dynamics
Recovering the scene’s state at a moment other than the one depicted, or the causal process connecting those moments. The media is one slice; the answer lies behind it, ahead of it, or in the mechanism driving it.
Retrodiction
Inferring what already happened based on visual traces.
Mechanistic Causality
Identifying the immediate physical mechanism driving an ongoing event. This is about understanding the “how” and “why” of current motion or state changes, not predicting the future.
Extrapolation
Predicting the immediate kinematic trajectory or long-term systemic outcome of the current scene.
Physical & Spatial Logic
Recovering unshown structure of the physical scene: what exists outside view, what properties objects hold beyond their appearance, and how the layout resolves from a coordinate the camera never occupied.
Hidden & Invisible
Reasoning about entities or physical properties that are not directly visible in a scene. Deducing occluded, off-screen, or internal objects, non-visual attributes (mass, friction), and invisible forces.
Affordance & Feasibility
Assessing the physical compatibility between objects or between agents and their environment.
Spatial “Alien Viewpoint”
Mentally re-projecting the 3D scene to calculate line-of-sight or reachability from a different physical coordinate.
Appendix
Table 4: Taxonomy Design for HSS along with descriptions.
Vendor
Model
Endpoint
API
Effort
Out
Native vid.
Frames
OpenAI
GPT-6-astra
openai/gpt-6-astra
Responses
max
128,000
no
500
GPT-6.1-sol
openai/gpt-6.1-sol
Responses
max
128,000
no
500
GPT-6-sol
openai/gpt-6-sol
Responses
max
128,000
no
500
GPT-5.6-sol
openai/gpt-5.6-sol
Responses
max
128,000
no
500
GPT-5.6-terra
openai/gpt-5.6-terra
Responses
max
128,000
no
500
GPT-5.6-luna
openai/gpt-5.6-luna
Responses
max
128,000
no
500
Appendix
Table 5: Model configurations for the main experiments. All 25 models of Table 1 , grouped by vendor in the same order, plus the three judges of Section 3.4 . Endpoint is the identifier sent to the proxy; API is the request shape (OpenAI Chat Completions, OpenAI Responses, or Anthropic Messages). Effort is the reasoning level used for the main table. Out is the requested output-token ceiling, which thinking tokens count against. Native vid. is whether the clip is sent as a video block rather than as sampled frames. Frames is the per-request frame cap that binds for that deployment; sampling is 2 fps for all models, thinned uniformly to fit the cap.
Family
Cause
Count
Share
Image
Video
Perception
Misidentified an object, person, or role
1,702
19.9
17.7
22.2
Misread orientation or facing
1,031
12.0
18.2
5.1
Misjudged size, distance, or fit
1,048
12.2
15.8
8.2
Read a 2D overlap as 3D alignment
776
9.1
14.6
2.9
Latent
Misread an invisible quantity from its trace
634
7.4
6.9
8.0
Misread motion or sequence
678
7.9
3.8
12.4
Appendix
Table 6: Primary failure causes, pooled over 24 models ( n=8,573 ). Share is the percentage of all failures; the last two columns split that share by media type. Perception and latent causes together account for 94% of failures.
Family
Cause
Single
Agentic
Δ %
Fixed %
Perception
Misidentified an object, person, or role
181
147
−19
35
Perception
Misread orientation or facing
117
93
−21
40
Perception
Misjudged size, distance, or fit
137
105
−23
39
Perception
Read a 2D overlap as 3D alignment
102
101
−1
24
Latent
Misread an invisible quantity from its trace
71
55
−23
37
Latent
Misread motion or sequence
83
52
−37
33
Appendix
Table 7: How the agent loop shifts the failure distribution , four harness/base pairs pooled on 387 paired tasks, first attempts. Fixed is the percentage of single-call failures with that cause that the harness repairs.
Subdomain
n
Single
Agentic
Δ
Patterns & pareidolia
68
33.8
64.7
+30.9
Retrodiction
152
34.9
55.3
+20.4
Theory of mind
136
32.4
48.5
+16.2
Social role, norm & power
168
32.1
47.6
+15.5
Affordance & feasibility
140
43.6
55.0
+11.4
Change & consequence
224
40.6
49.1
+8.5
Appendix
Table 8: Where the agent loop pays, by subdomain. Pass@1 (%) for the four harness/base pairs pooled on the 387 paired tasks of Table 7 , first attempts. n counts pair × task observations, so each subdomain appears four times. Re-inspection helps most where the evidence is in the frame and least where the answer turns on depth or on an outcome the media never shows.
Figure 9: One representative failure per subdomain (1 of 6): the two subdomains of physical & spatial logic that turn on what lies outside the frame: hidden & invisible, and the alien viewpoint. Every task shown was missed by GPT-6-astra, Gemini-3.8-flash and Claude-Opus-5.5 on all three attempts; the model line is the judge’s condensed reading of what all three said. All eleven cases are image tasks, so each card shows exactly the media the models were served.
Figure 10: One representative failure per subdomain (2 of 6): the two that turn on fit and reach: affordance & feasibility, and spatial reachability. Selection and panel layout as in Figure 9 .
Figure 11: One representative failure per subdomain (3 of 6): the forward and backward temporal subdomains: extrapolation and retrodiction. Selection and panel layout as in Figure 9 .
Figure 12: One representative failure per subdomain (4 of 6): mechanistic causality and theory of mind. Selection and panel layout as in Figure 9 .
Figure 13: One representative failure per subdomain (5 of 6): social role, norm & power, and change & consequence. Selection and panel layout as in Figure 9 .
Figure 14: One representative failure per subdomain (6 of 6): patterns & pareidolia. Selection and panel layout as in Figure 9 .
Harness
Base model
Raw
Agent
Δ
Win
Loss
p
Think
Blank
Claude Code
Claude-Opus-5
30.8
51.3
+20.5
112
30
<10−4
43,764
0
Claude Code
Claude-Fable-5.1
41.2
57.7
+16.5
96
30
<10−4
42,237
1
Codex
GPT-5.6-sol
28.4
35.8
+7.5
63
33
0.003
6,541
4
Codex
GPT-6-astra
54.4
59.3
+4.9
53
28
0.007
2,040
0
Appendix
Table 9: All agentic harness/base pairs , paired on 388 tasks. Blank counts tasks where the harness returned no answer, scored as zeros; p is two-sided McNemar.
Subdomain
Reasoning level, lowest → highest
Δ
GPT-6-astra
Hidden & Invisible
48
50
52
57
60
+12
Affordance & Feasibility
39
41
44
49
56
+17
Spatial “Alien Viewpoint”
28
30
40
46
44
+16
Spatial Reachability
22
22
26
22
22
+0
Extrapolation
41
45
47
49
53
+12
Appendix
Table 10: Accuracy by subdomain and reasoning effort (%), all 522 tasks. Columns are each deployment’s levels in order — five for GPT-6-astra ( low – max ), four for Claude-Opus-5.5, three for Gemini-3.8-flash — and Δ is last minus first.
Model
10
50
100
200
500
Δ
best
GPT-6-astra
33.8
41.9
45.3
47.0
51.7
+17.9
500
Claude-Opus-5.5
30.3
32.1
32.5
37.2
36.3
+6.0
200
GPT-6-sol
23.1
27.8
29.1
26.1
31.2
+8.1
500
Claude-Opus-5
18.8
19.7
28.6
27.4
27.4
+8.5
100
Appendix
Table 11: Video frame budget. Pass@1 (%) on the 234 video tasks at each per-request frame cap, sampled at 2 fps and thinned uniformly to fit. Δ is 500 minus 10 frames; best is the budget at which each model peaks.
Judge
Model
Claude-Opus-5
GPT-5.6-terra
Gemini-3.6-flash
Spread
GPT-6-astra (max)
53.8
50.8
54.2
3.4
Gemini-3.8-flash (high)
42.7
40.8
43.5
2.7
Claude-Fable-5 (max)
33.8
30.2
33.0
3.6
Claude-Opus-5 (max)
31.5
30.0
31.7
1.7
GLM-5.3-flash (max)
26.9
25.2
27.1
1.9
Appendix
Table 12: Cross-judge alignment. Pass@1 (%) for the same five models under three judges from three vendors, single attempt per task. The rightmost column is the largest disagreement between any two judges about a model. Every pairwise rank correlation is ρ=1.0 : the three judges disagree about levels, never about order. Agreement statistics for the same runs are plotted in Figure 6 .
Recent multimodal large language models (MLLMs) achieve strong performance on visual reasoning benchmarks, yet it remains unclear to what extent such performance reflects reasoning directly grounded in visual evidence. We introduce VisReason, a benchmark for vision-centric reasoning in everyday scenarios where perception and inference are tightly coupled. VisReason contains 1,505 questions across 10 categories spanning perceptual, structural, and conceptual reasoning. Our evaluation shows that VisReason poses a qualitatively different challenge from existing benchmarks, exposing substantial gaps between humans and current MLLMs and revealing limited benefits from test-time reasoning strategies. VisReason offers a focused diagnostic for evaluating vision-centric reasoning beyond language.
Longteng Guo, Yifan Wang, Pengkang Huo +4
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences
While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, lagging behind 6-year-old humans and falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs still lack fundamental visual primitives. Progress in BabyVision represents a step toward human-level visual perception and reasoning capabilities. We also explore solving visual reasoning with generation models by proposing BabyVision-Gen and automatic evaluation toolkit. Our code and benchmark data are released at https://github.com/UniPat-AI/BabyVision for reproduction.
Liang Chen, Weichu Xie, Yiyan Liang +27
1UniPat AI · 6Peking University · 3Alibaba Group +8
Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they attain high scores on many popular visual benchmarks, with headroom rapidly eroded by model progress. This creates a need for difficult benchmarks that remain relevant for longer. We introduce ZeroBench - a lightweight visual reasoning benchmark curated using adversarial filtering to be "impossible" for frontier LMMs at its original release, with initial SotA scores of 0% pass@1 and pass^5. We track progress on ZeroBench over the subsequent year, observing SotA reaching 6% pass^5 and 19% pass@5, indicating the potential longevity of the benchmark. We evaluate 46 LMMs on ZeroBench, compare performance to a human baseline, analyse strengths and weaknesses, chart a year of progress in visual capabilities, and publicly release ZeroBench at https://zerobench.github.io.
Jonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma +31