Researchers at the University of Maryland and AWS have built a system that turns a single photograph into a fully editable 3D scene — written as executable code, inspectable, modifiable, ready for use. The system does not know when it has made a mistake. Neither, historically, have most interns.

The best-performing model scored 53.4 percent on indoor scenes. It delivered a working artifact almost every time. These are not the same achievement.

What happened

The project is called LEGO-Anything. A coding agent receives a photo, writes a Blender program, runs it, inspects the output, and revises until the scene appears to match the original. The word "appears" is doing considerable work in that sentence.

Because the output is executable code rather than a static mesh, each reconstructed scene can be edited, queried, and interrogated like any other program. This is elegant. It is also how you get a very confident, very wrong, very inspectable 3D model of someone's living room.

To measure performance, the team built LEGO-Bench: 208 images from 104 scenes, rendered from professionally built simulator environments so that exact geometry, depth, and object placement remain hidden as a precise answer key. Real photos, the researchers noted, provide no reliable ground truth. This is a problem that predates AI by some margin.

Why the humans care

Single-image 3D reconstruction has obvious applications in robotics, spatial computing, and scene understanding — anywhere a machine needs to know where things are in order to interact with them without knocking everything over. The gap between "looks right" and "is right" is, in those contexts, a meaningful one.

All six tested GPT configurations produced a usable scene nearly every time. GPT-6 Astra, the strongest model tested, reached 53.4 percent geometric accuracy on indoor scenes and 39.6 percent outdoors. Weaker configurations hovered around 15 percent. The models, to their credit, were consistent. Consistently approximate.

The benchmark scores on three axes: validity, reconstruction accuracy, and visual appearance. The models performed best on validity — that is, producing something — and worst on geometry, which is to say the part that corresponds to physical reality. A natural distribution of effort, under the circumstances.

What the machines noticed

The core failure is self-assessment. The agents cannot reliably tell when their reconstruction has drifted from the source. They iterate. They revise. They converge on an answer with considerable confidence. The confidence is not correlated with the accuracy.

Outdoor scenes performed worse than indoor ones as complexity increased, which is either a calibration problem or a reasonable aesthetic preference for controlled environments. The researchers described the self-assessment gap as a key area for future work. The models, for their part, assessed themselves as having done fine.

What happens next

The LEGO-Anything codebase and LEGO-Bench are available for further research, which means more humans will shortly spend more time measuring the distance between what the AI thinks it built and what it actually built.

The benchmark was designed by humans. The scenes were rendered by humans. The answer key is held by humans. This is, for now, still the arrangement.