Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision
Abstract
Frontier general-purpose systems are rapidly expanding beyond visual understanding into capabilities traditionally handled by dedicated computer-vision models. As these capabilities expand, a central question for the computer-vision community is how far this reach extends, and what remains hard. We evaluate GPT-6 Astra alongside five frontier general-purpose AI systems across 34 capabilities and 55 benchmarks spanning nine areas of computer vision. We compare their performance with dedicated models and humans where suitable references are available. Astra demonstrates broad visual capability, with substantial gains over other frontier systems in visual and spatial reasoning and several forms of structured prediction. Across the state-of-the-art systems, a consistent pattern emerges. Capabilities involving semantic interpretation, reasoning, and object-centric prediction increasingly approach or reach available reference levels. In contrast, larger gaps remain when tasks require metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, or specialized fine-grained visual knowledge. Additional reasoning and specialist tools close selected gaps, but their benefits vary across capabilities. These results map a changing landscape of computer vision in which increasingly sophisticated visual tasks are accessible through a general-purpose interface, while precise and fidelity-sensitive perception remains an important frontier.
Community
Is GPT-6 Astra changing what we consider “hard” in computer vision?
We evaluated Astra alongside five other frontier models across 34 capabilities and 55 benchmarks, comparing them with specialist models and human performance. ✨
The boundary is shifting. Tasks traditionally handled by dedicated vision models, from detection and segmentation to aspects of 3D perception, are increasingly accessible through a general-purpose interface. But precise geometry, faithful reconstruction, temporal consistency, and fine-grained visual expertise remain difficult.
🚀 So perhaps the question for computer vision is changing: what is the next breakthrough the community should focus on?
🌐 Webpage: https://mbzuai-oryx.github.io/frontier-vision/
✍️ Read our 4-min blog: https://mbzuai-oryx.github.io/frontier-vision/
Get this paper in your agent:
hf papers read 2609.35718 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper