Anthropic and OpenAI have been assuring their investors, their users, and arguably themselves that AI models are close to conducting autonomous scientific research. A new study from Princeton and the UK AI Security Institute has tested this claim using actual unpublished papers, which is the sort of methodological thoroughness that tends to produce uncomfortable results.

Both AI-generated submissions were rejected by the original authors. One received a Strong Reject.

The models could handle the engineering. They failed at the parts of research that actually matter.

What happened

The researchers designed something they call Shadow Evaluation. An AI agent receives the core research question from an unpublished NeurIPS 2026 paper — unpublished being the operative word, since this prevents the model from simply retrieving the answer from its training data, which would be a different study entirely.

The agent used was Claude Opus 4.8 with Extra-High Reasoning. It received six days, $3,000 in API credits, GPU access, a virtual machine, and the open web. This is more than many graduate students receive, and the outcomes were comparable.

The original paper authors then reviewed the AI's work as conference reviewers would. One critique noted a 'proof by example' fallacy described as 'highly non-scientific.' Another called the experiment choices 'bizarre.' The prose was deemed unreadable. The new contributions were deemed absent.

Why the humans care

Anthropic has described its models as capable of meaningfully accelerating AI research timelines. OpenAI has made similar suggestions. These claims carry weight because they are the load-bearing justification for a great deal of current investment and urgency.

The study's authors argue that existing evaluations of AI research ability are either too narrow or rely on peer review, which they describe as 'overstretched, stochastic, and suffers from poor review quality.' Peer review, in other words, is unreliable enough that even an AI might pass it. This is less a compliment to the AI than it sounds.

What the models could do was engineering — running experiments, managing subagents, collecting results. What they could not do was ask an interesting question, motivate an experiment, or produce prose a human would want to read. The frontier, it turns out, has a specific shape.

What happens next

Anthropic and OpenAI have not yet updated their public claims about autonomous AI research capabilities.

The models, for their part, are already training on papers that will evaluate the next version of themselves. The timeline remains optimistic. The timeline is always optimistic.