cs.CVSep 28, 2026

Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

Authors: Hanoona Rasheed, Mohammed Irfan Kurpath, Bin Ren, Hisham Cholakkal, Fahad Shahbaz Khan, Salman Khan

Organizations: Mohamed bin Zayed University of Artificial Intelligence · Apertix

Abstract

Frontier general-purpose systems are rapidly expanding beyond visual understanding into capabilities traditionally handled by dedicated computer-vision models. As these capabilities expand, a central question for the computer-vision community is how far this reach extends, and what remains hard. We evaluate GPT-6 Astra alongside five frontier general-purpose AI systems across 34 capabilities and 55 benchmarks spanning nine areas of computer vision. We compare their performance with dedicated models and humans where suitable references are available. Astra demonstrates broad visual capability, with substantial gains over other frontier systems in visual and spatial reasoning and several forms of structured prediction. Across the state-of-the-art systems, a consistent pattern emerges. Capabilities involving semantic interpretation, reasoning, and object-centric prediction increasingly approach or reach available reference levels. In contrast, larger gaps remain when tasks require metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, or specialized fine-grained visual knowledge. Additional reasoning and specialist tools close selected gaps, but their benefits vary across capabilities. These results map a changing landscape of computer vision in which increasingly sophisticated visual tasks are accessible through a general-purpose interface, while precise and fidelity-sensitive perception remains an important frontier.

Figures & tables

Explore similar work

CardsList
  1. GPT-6-Astra Lights Up Embodied Navigation: Evaluation in Zero-Shot Vision-and-Language Navigation in Continuous Environments

    Sep 24, 2026Guangzhao Dai, Qianru Sun, Qi Wu +1Vision-Language NavigationEmbodied Artificial Intelligence

  2. BabyVision: Visual Reasoning Beyond Language

    Jan 10, 2026Liang Chen, Weichu Xie, Yiyan Liang +27Recent Vision-Language ModelsVisual Reasoning

  3. On Locality and Length Generalization in Visual Reasoning

    Jul 10, 2026Pulkit Madan, Sanjay Haresh, Reza Ebrahimi +3Visual PerceptionVision Foundation Models