OpenAI's GPT-6 Astra can now tell you exactly where you screwed up your IKEA shelf

OpenAI's GPT-6 Astra has demonstrated a notable leap in visual reasoning, correctly identifying assembly mistakes in IKEA furniture pieces roughly 80 percent of the time when shown a photograph. That figure represents a dramatic improvement over where things stood in November 2025, when the best available model could only manage 28 percent accuracy on the same task. The benchmark, tracked by Epoch AI, has become a useful proxy for how well multimodal models handle real-world spatial problem-solving.
The IKEA assembly task is a meaningful test because it demands more than simple object recognition. A model must understand 3D structure from a 2D image, cross-reference it against a known correct configuration, and pinpoint a specific discrepancy - such as a bracket facing the wrong direction or a panel inserted out of sequence. These are the kinds of fine-grained spatial comparisons that earlier vision-language models consistently struggled with.
Despite the strong accuracy numbers, Epoch AI notes that the system is not yet fast enough to serve as a real-time assembly assistant - the kind of tool that could watch over your shoulder via a phone camera and flag mistakes as you make them. Inference speed and latency remain limiting factors, meaning the practical use case is still closer to a post-hoc error checker than an interactive guide. That said, the gap between current performance and the threshold needed for live assistance appears to be narrowing.
The broader significance of this benchmark is what it suggests about the trajectory of multimodal AI. Tasks requiring physical-world spatial understanding - assembly, repair, navigation - have historically been a weak point for large language models with vision capabilities. An improvement of this scale in under a year points to continued rapid development in that area, with practical applications in manufacturing, home improvement, and accessibility potentially within closer reach than previously assumed.
