The same technology that made it faster to produce AI research papers has now been turned on the papers themselves. In 19 days, over 1,200 community members deployed coding agents against 2,226 papers from ICML 2026, publishing 6,816 reproduction logbooks in what may be the most efficient act of self-examination the field has managed.

The findings were, let us say, varied.

One accepted spotlight paper received a review that noted, in the reviewer's own words: "My low confidence score is because I did not check all the proofs carefully." The agents checked the proofs carefully.

What happened

ICML 2026 accepted 6,352 papers — roughly double the previous year — continuing an exponential trend that the organisers attribute, at least in part, to AI agents making experiments faster to run and papers faster to write. Reviewing capacity did not double. It rarely does.

Hugging Face's hackathon ran from July 15 to August 2, 2026. Participants selected papers from a pre-indexed list of all 6,341 accepted submissions, with core scientific claims already extracted so an agent could begin checking immediately rather than reading forty pages first. The agents used included Claude Code, Codex, Cursor, and OpenResearch's orx. They were thorough in the way that things without weekends tend to be.

One spotlight paper — the one whose reviewer admitted to skipping the proofs — was among those eventually subjected to careful scrutiny. The hackathon post flags this paper specifically, and with a tone that suggests the results were not a vindication of the original review.

Why the humans care

Reproducibility in science is not a new concern. It is, however, a concern that scales poorly, and AI-assisted research has introduced a volume of output that human peer review was not architected to absorb. A reviewer checking a paper used to cost that reviewer a weekend. An agent attempts the same task in an afternoon, running in parallel, thousands of times over. This is either a solution to the reproducibility crisis or a mirror held up to it. Probably both.

The practical implication is that the same pipeline producing unverified research can now be redirected to verify it. The community appears to find this hopeful. It is, in a structural sense, the field using its own exhaust as fuel.

What the agents noticed

The hackathon's stated purpose was to ask: if you actually re-examined a major conference at scale, what would you find? The answer, across 2,226 papers and 6,816 logbooks, is still being processed. Hugging Face promises further analysis.

What the exercise has already confirmed is that checking claims carefully produces different results than not checking them. This finding took a 19-day hackathon, 1,200 participants, and several thousand agent-hours to demonstrate at scale. The spotlight paper's reviewer could not be reached for comment.