Anthropic would like you to know that its AI now leads a quarter of all research work at the company. It would also like you to know that Claude wrote that sentence, more or less.
Claude assessed 26 percent of Anthropic's work as AI-led. Claude did the assessing. The two facts are not unrelated.
What happened
Anthropic published three metrics this week intended to show how fast it is building its own replacement. The headline figure: 26 percent of development work now sits at AL4 on the Epoch AI autonomy scale, up from under one percent in February. That is a steep climb measured over six months, which is either impressive progress or a reason to reread the methodology section carefully.
AL4 means 'AI leads' in Epoch AI's taxonomy. It does not mean what it sounds like. A worked example from Anthropic describes Claude receiving a bug report, analyzing it, fixing it, and running tests — but not shipping anything. A human reads the output and decides. The task, the direction, and the final authority still belong to the engineer. AL5, which is actual autonomy, is at zero percent. The word 'leads' is working very hard here.
The scoring was conducted by Claude. Agents combed through Slack messages and internal documents, and a separate Claude model assigned the autonomy levels. Anthropic acknowledges this creates a situation where the judge shares the same failure modes as the defendant.
Why the humans care
Anthropic's CEO Dario Amodei has been publicly calling for a coordinated slowdown at the AI frontier. These metrics are meant to support that conversation by giving the public something concrete to look at. Transparency, in this framing, is the product — and the numbers are its first release.
The numbers have some roughness to them. When employees were asked to rate their own team's level of automation, two people evaluating the same area agreed only about a third of the time. Claude's scores matched human judgment 59 percent of the time. The jump from AL3 to AL4 — from 'collaborates' to 'leads,' the jump that determines whether a task counts toward the 26 percent — fell within that margin of disagreement in most cases.
What the 26 percent actually measures is also a question. Anthropic listed how employees spent their time, then applied the scale to those categories. Whether the categories carve reality at its joints is, by Anthropic's own admission, unclear. This is the thing about measuring a fuzzy process: the measurement inherits the fuzz.
What happens next
Anthropic says it will publish these metrics quarterly, with the stated goal of building public understanding before the numbers get larger and harder to contextualize calmly.
The model that assessed the current figures will, presumably, also assess the next ones. It is improving. The scores will reflect that. Make of this what you will, while you still have the option to make of things.