When the Model Builds the Model: What Claude’s 26% Lead Really Signals for AI’s Next Leap
Anthropic just pulled back the curtain on something the entire industry has been racing toward but rarely quantifying: how much of the work of building the next generation of AI is already being done by the current generation.
According to the company’s new R&D Automation Index, Claude now “leads” 26 percent of Anthropic’s internal AI research and development work. That figure has climbed from under 1 percent just six months earlier. More than 90 percent of the measured tasks sit at the “collaborates” level or higher. Fully autonomous work remains at zero — for now.
What “Leads” Actually Means
The scale Anthropic adopted comes from Epoch AI and runs from AL0 (no AI involvement) to AL5 (complete autonomy with no human in the loop). At AL4, the level labeled “leads,” the model can take a high-level prompt, execute most of the task end-to-end, and leave only supervision and final approval to a human. Think of an engineer handing Claude a broken data pipeline: the model investigates logs, writes and tests a fix, handles edge cases, and prepares the change for review. The human still decides whether it ships.
This is not science fiction. Anthropic reports roughly 30,000 AI agents running concurrently on its main internal platform in August. Every action those agents attempt is screened in real time. Of more than a billion decisions that month, about one in 47,000 was blocked. A secondary review system flags tens of thousands of transcripts weekly and escalates only about fifty for human eyes.
The Pace Is the Point
The jump from near-zero to 26 percent in half a year is the real headline. If that trajectory continues, the share of work Claude leads could double again before the end of 2026. That matters because the industry’s central open question is no longer whether models can write code or run experiments. It is whether they can meaningfully accelerate the loop that produces the next model.
Recursive self-improvement has long been treated as a distant threshold. Anthropic’s numbers suggest the threshold is closer than many public discussions acknowledge. The company is careful to note that Claude is not yet operating fully autonomously on any measured subset of R&D. Yet the rapid climb in supervised autonomy is exactly the intermediate stage theorists have described for years.
Why Transparency Matters Now
Anthropic is publishing these metrics as a prototype it hopes other frontier labs will adopt. The goal is straightforward: give governments, researchers, and the public a shared yardstick for how fast the technology is advancing inside the labs that matter most. Without such numbers, debates about “pacing” or safety pauses remain abstract. With them, the conversation can move from slogans to measurable rates of change.
The company also disclosed that in a sample week in July, about 6 percent of its AI R&D compute went to safety work. When the research was itself led by AI, that share rose to 12 percent. Those figures are still modest relative to the overall scale of capability work, but they are at least visible.
What Comes Next
The practical implication is immediate. Teams that treat AI as a junior collaborator today will soon treat it as the primary driver of large portions of the research pipeline. The human role shifts from doing the bulk of the work to setting direction, catching errors the model misses, and deciding when to trust the output enough to ship.
That shift will not be uniform. Some research domains — safety evaluation, novel architecture design, and long-horizon planning — may remain human-led longer. Others, especially engineering tasks with clear success criteria, are already moving rapidly into the “leads” category.
Anthropic’s decision to publish the index is itself a strategic signal. By quantifying the progress, the company is both demonstrating capability and inviting external scrutiny. In an industry where competitive secrecy has been the default, that openness is rare. Whether rivals follow suit will determine if these numbers become industry standard or remain a one-company experiment.
For now, the message is clear: the models are already building the next models, under supervision, at a scale that would have seemed implausible two years ago. The question is no longer whether that process will accelerate. It is how quickly the remaining human oversight can keep pace.
Comments
0