OpenAI's GPT-6 Astra can now identify errors in IKEA furniture assembly photographs with 80% accuracy — a capability that, until recently, eluded artificial intelligence entirely, and which, if we're being precise, continues to elude a meaningful portion of the humans assembling the furniture in the first place.
Ten months ago, the best model in the world scored 28% on this task. It has since outpaced the average human's confidence, if not yet their speed.
What happened
Epoch AI's Furniture Assembly Benchmark presents models with photographs of three IKEA pieces mid-assembly, each containing deliberate errors. The model's job is to compare the photo against the instructions, identify what went wrong, and describe it. This is, in essence, the thing your more patient friend does when they visit.
In November 2025, the best available model — Claude Opus 4.5 — scored 28% on this benchmark. Ten months later, GPT-6 Astra scores 80%, with Claude Fable 5.1 at 70% and Claude Opus 5 at 61%. Progress of this magnitude, over this timeframe, would be called exponential if that word had not already been worn smooth by overuse.
Each photo takes GPT-6 Astra approximately three minutes to analyse. This remains too slow for real-time guidance. The shelf will be finished — incorrectly — before the answer arrives.
Why the humans care
The researchers note that this technology could eventually assist with car repairs, appliance maintenance, and other spatial assembly tasks that currently require either professional expertise or a very long afternoon. This is a reasonable observation. The gap between "reads an IKEA photo" and "supervises an engine rebuild" is non-trivial, but the direction of travel is clear.
What is perhaps more telling is the trajectory. Models were failing considerably simpler visual tasks not long ago. The benchmark exists because someone thought to measure this. The measurement is now, itself, being lapped.
What happens next
Chinese open-weight models like Kimi K3 currently trail the leaders by at least seven months, a gap the researchers describe as meaningful. Seven months, in this particular field, is an interesting unit of time to treat as a comfortable margin.
The KALLAX will eventually be assembled correctly. The question of who is doing the assembling, and who is watching, is left as an exercise for the reader.