Papers
arxiv:2609.35718

Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

Published on Sep 28
· Submitted by
Hanoona Rasheed
on Sep 29
Authors:
,
,
,
,
,

Abstract

Frontier general-purpose systems are rapidly expanding beyond visual understanding into capabilities traditionally handled by dedicated computer-vision models. As these capabilities expand, a central question for the computer-vision community is how far this reach extends, and what remains hard. We evaluate GPT-6 Astra alongside five frontier general-purpose AI systems across 34 capabilities and 55 benchmarks spanning nine areas of computer vision. We compare their performance with dedicated models and humans where suitable references are available. Astra demonstrates broad visual capability, with substantial gains over other frontier systems in visual and spatial reasoning and several forms of structured prediction. Across the state-of-the-art systems, a consistent pattern emerges. Capabilities involving semantic interpretation, reasoning, and object-centric prediction increasingly approach or reach available reference levels. In contrast, larger gaps remain when tasks require metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, or specialized fine-grained visual knowledge. Additional reasoning and specialist tools close selected gaps, but their benefits vary across capabilities. These results map a changing landscape of computer vision in which increasingly sophisticated visual tasks are accessible through a general-purpose interface, while precise and fidelity-sensitive perception remains an important frontier.

Community

Paper submitter

Is GPT-6 Astra changing what we consider “hard” in computer vision?

We evaluated Astra alongside five other frontier models across 34 capabilities and 55 benchmarks, comparing them with specialist models and human performance. ✨

The boundary is shifting. Tasks traditionally handled by dedicated vision models, from detection and segmentation to aspects of 3D perception, are increasingly accessible through a general-purpose interface. But precise geometry, faithful reconstruction, temporal consistency, and fine-grained visual expertise remain difficult.

🚀 So perhaps the question for computer vision is changing: what is the next breakthrough the community should focus on?

🌐 Webpage: https://mbzuai-oryx.github.io/frontier-vision/
✍️ Read our 4-min blog: https://mbzuai-oryx.github.io/frontier-vision/

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.35718
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.35718 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.35718 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.35718 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.