BlogSeptember 20264 min read

Hard Vision, Easy Vision What GPT-6 Astra reveals across computer vision.

Successive generations of frontier models support a growing range of computer vision tasks, including tasks traditionally handled by specialist models. For computer vision researchers, this raises a broader question: how much of computer vision is now accessible through a general-purpose interface, and where do dedicated models remain necessary? Existing evaluations provide substantial evidence, but their findings are spread across benchmarks, papers, and model comparisons. Bringing these results together can help clarify which capabilities are established across models, where new capabilities are emerging, and which challenges remain.

The study. We organized our study around 34 capabilities across nine broad areas of computer vision, drawing on 55 benchmarks. We compare six frontier models with one another to identify shared strengths and differences; with specialist models, built specifically for individual tasks, to assess their performance on traditionally specialized tasks; and with reported human performance to place these results in context.

6
frontier models
34
capabilities
55
benchmarks
9
areas of vision

Frontier visual capability. Our evaluations show that the visual capabilities frontier models extend beyond image recognition and description to object detection, image and video segmentation, and 3D perception, including depth estimation and reconstruction. Their capabilities also extend to tasks requiring domain expertise, such as medical-image understanding and grounding, and remote sensing, as well as embodied understanding and navigation, image generation, and editing. These results indicate that a broad range of tasks traditionally handled by dedicated vision models is becoming accessible through general-purpose models, although performance varies across tasks.

Ground truth GPT-6 Astra
Aerial view of a parking lot with detection boxes on every car
2D detection · 146 densely packed “cars”
Tray of steamed buns with detection boxes
2D detection · crowded “buns”
Rabbit tattoo on a woman's back
Segmentation · “the rabbit on the woman's back”
Predicted depth map
Ground-truth depth map
Ground truthGPT-6 Astra
Depth estimation · layout captured, fine detail smoothed
3D visual grounding of a blackboard in a point cloud
3D visual grounding · “the thin, 4-meter blackboard”
3D bounding boxes on a street scene
3D object detection · coherent 3D boxes

Established strengths. Some capabilities are already becoming established strengths across frontier models. Recognition, visual reading, and knowledge-based reasoning stand out for their consistency across the six models we evaluated. These models combine visual understanding with world knowledge to read documents and charts, interpret scientific content, and solve visual mathematics problems. On the evaluated benchmarks, OCR approaches human performance, while all six models surpass the reported human reference in document, chart, and infographic understanding, as well as visual mathematical reasoning.

Consistent across all six models
Scores of the six frontier models, with the reported human and specialist references as dashed lines.
GPT-6 AstraOther five modelsHumanSpecialist

Emerging capabilities. Other capabilities show larger differences across models. Astra performs particularly strongly on several of these tasks. On the evaluated benchmarks, it surpasses the specialist references in object detection and image segmentation. It also exceeds the next-best frontier model in visual logical reasoning (+13.6 points), pose estimation (+22.7), video segmentation (+26.6), and 3D medical grounding (+27.9). In 2D spatial reasoning, it reaches the reported human performance. These strengths are not yet shared consistently across frontier models, but they expand our picture of what is becoming possible.

Strong in one model, not yet across the frontier
Scores of the six frontier models against the specialist or human reference.
GPT-6 AstraOther five modelsSpecialistHuman

Semantic competence versus perceptual precision. Across these results, a recurring distinction emerges between semantic competence and perceptual precision. Frontier models show stronger progress in interpreting visual content and reasoning about it. Larger gaps remain when tasks require accurate geometric measurements, faithful preservation of fine details, or precise predictions maintained across video frames.

A recurring distinction between semantic competence and perceptual precision.

This distinction is also visible across related tasks. In 3D perception, models can locate an object from a description, yet still struggle to estimate depth accurately or reconstruct the surrounding scene. In image generation and editing, they can produce visually appealing results, but tasks like restoration are more demanding because they require faithful recovery of the original details. In video segmentation, although Astra leads the next-best frontier model by 26.6 points, it still faces significant challenges with precise boundaries, small structures, and consistency across frames. In medical imaging, similarly strong understanding and reasoning can coexist with gaps in specialized visual expertise. Models can draw on medical knowledge to interpret an image and even locate fine structures such as individual nuclei, yet struggle to distinguish nucleus types, which requires recognizing subtle visual differences, particularly in rare or less familiar categories.

Predicted video segmentation
Ground-truth video segmentation
Ground truthGPT-6 Astra
Video segmentation · instances are separated; boundaries and small parts drift
Ground-truth image
Restored image
GPT-6 Astra + image modelGround truth
Restoration · plausible, but not faithful: pins, lanes and banner text change
Zoomed pathology crop with GPT-6 Astra nucleus boundaries
Zoomed pathology crop with ground-truth nucleus boundaries
Ground truthGPT-6 Astra
Pathology · nuclei are easy to locate; telling their types apart is the hard part

This blog only scratches the surface. In the paper, we dig into four questions: how broad frontier visual capability really is, where the models agree and where they differ, where specialist-level performance is emerging, and whether a bit more reasoning or a few tools can close the remaining gaps (and what that costs). We hope you enjoy the read!