Period ending 2026-09-21
30 new papers
A weekly snapshot of new work published in Large Vision Language Models.
Twelve weeks of publication activity for this topic as it is defined today.
Weekly history
What was published in this field, kept on the site without email delivery.
Period ending 2026-09-21
A weekly snapshot of new work published in Large Vision Language Models.
Period ending 2026-09-14
A weekly snapshot of new work published in Large Vision Language Models.
Period ending 2026-09-07
A weekly snapshot of new work published in Large Vision Language Models.
Inside this field
Within Large Vision Language Models
Within Large Vision Language Models
Within Large Vision Language Models
Within Large Vision Language Models
Within Large Vision Language Models
Within Large Vision Language Models
Within Large Vision Language Models
1,309 papers
Perception-then-Reasoning'' paradigm degenerates perception into a passive, one-off feature encoding process, rendering it functionally equivalent to Reasoning-in-Text-Space'', where task-critical spatial signals are collapsed before reasoning begins. We substantiate this claim with the Turing Eye Test (TET): tasks that must be resolved in \emph{visual space} and are hard to verbalize; results show text-only reasoning cannot remedy these perceptual failures. Our findings suggest rethinking the architectural divide: shifting from reasoning \textit{about} perception to reasoning \textit{within} perception. This facilitates actively reasoning-driven perception that operates directly on pixel-level visual representations, rather than within a collapsed textual space.