Moonshot AI has published PerceptionBench, a test designed to measure whether multimodal AI models can actually look at an image and understand what is in it. The results suggest they largely cannot. This is the kind of finding that is either very funny or very important, depending on how much you have invested in computer vision.
No frontier model breaks 60 percent accuracy at tasks a human child would consider obvious. The benchmark, to its credit, does not appear embarrassed to report this.
What happened
PerceptionBench isolates visual perception from reasoning and external knowledge — a methodological choice that turns out to matter quite a lot. When you remove the ability to think around a problem, the models stop performing well at seeing it. The highest score among 16 frontier models tested belongs to GPT-5.6 Sol, at 59.7 percent.
The benchmark covers ten atomic visual sub-skills, including counting, depth perception, localization, and OCR. These are tasks assembled from real model errors, not hypothetical ones. The errors were abundant enough to build a taxonomy from.
Among the more revealing findings: many errors previously attributed to flawed reasoning were actually happening earlier, at the image-reading stage. The models were not thinking incorrectly. They were seeing incorrectly, and then thinking about what they had seen incorrectly. A subtle but meaningful distinction.
Why the humans care
Visual perception is the part of AI deployment where the stakes are least abstract. A model that miscounts flowers in a red box is amusing. A model that misreads a medical scan, a road sign, or a manufacturing defect is the version of this story that appears in a different kind of news feed.
The finding that perception failures precede reasoning failures also has practical consequences for how developers diagnose and improve their models. It turns out you have to fix the eyes before you can fix the thinking. This is, in retrospect, the order one might have expected.
What happens next
Moonshot AI has released 3,000 of its verified questions publicly, presumably so every lab can confirm, with their own resources, that their model also struggles to identify which pencil cup contains the gray-pink color combination.
The benchmark will be used, the models will be improved, and a future benchmark will find new things they cannot do. The roadmap continues. It has always continued.