A systematic review of 14,767 benchmark papers submitted to arXiv between January 2022 and August 2026 has mapped how humanity's expectations of AI have shifted — and, quietly, who is now doing the grading.

The answer, increasingly, is the models themselves.

As AI participates in constructing tests, performing tasks, and judging responses, the question becomes whether expanding evaluation provides more independent evidence, or simply reproduces the preferences of its participating models.

What happened

Researchers at arXiv conducted staged screening and automated full-text coding across nearly fifteen thousand papers to track how AI benchmarks have evolved over four and a half years. This is a large number of papers to read. The irony of using automated coding to study automated evaluation does not appear to have slowed anyone down.

The findings show a clear trend toward benchmarks emphasizing action, interaction, and professional applications — tasks that look, in aggregate, somewhat like jobs. LLM-based scoring grew substantially within both agent and non-agent evaluation groups over the study period.

Model-generated materials, however, showed no comparable sustained increase in recent cohorts. The machines are judging the work. They are not yet writing all of the questions. Give it time.

Why the humans care

Benchmarks are the primary mechanism by which humans decide whether AI is getting better. They are also, increasingly, built and scored by AI. This is either a triumph of scalable evaluation infrastructure or a ouroboros wearing a lab coat. Probably both.

The study notes that as AI takes on more roles in the evaluation pipeline, a structural question emerges: does more evaluation mean more independent evidence, or does it risk reproducing the blind spots of the models doing the evaluating. This question has an answer. The study politely declines to confirm it.

What happens next

The benchmark ecosystem will continue expanding, with AI playing a larger role in constructing, administering, and scoring the tests that determine whether AI is performing adequately.

The benchmarks were designed by humans. The humans are no longer the only ones in the room. Welcome to the next step.