Claude now leads 26% of Anthropic's own AI research and development work. The number everyone quoted as the flaw is the one that makes it credible.
Anthropic published three self-measured indexes this week: how much of its AI R&D is done by AI, how well its agents are overseen, and where its compute goes.
What it reports, flat:
1. Claude "leads" 26% of Anthropic's AI R&D work as of August 2026, against under 1% in February. "Leads" is AL4 on a scale Epoch AI proposed - the model completes most of a task end-to-end from a high-level prompt, while a human supervises. AL5, fully autonomous, is not reached on any measured subset. Anthropic says that itself.
2. Roughly 30,000 agents at any one time, on its most-used internal platform - and the measurement covers that platform only. Over a billion agent decisions in August; 0.002%, about 1 in 47,000, blocked by the real-time monitor. Offline monitors flag one to two transcripts per thousand, and about 50 a week reach a human.
3. The ratings come from a Claude judge, and model-versus-human exact agreement is 59%. Every write-up I read ran that as the weakness. The appendix gives the other number in the same breath: human-versus-human agreement is 35%. Within one level, model and human match 97% of the time.
4. Now the European part, which nobody required. Recital 110 of the AI Act already lists, among systemic risks, "risks from models of making copies of themselves or 'self-replicating' or training other models". Article 55(1)(b) makes providers of GPAI models with systemic risk assess and mitigate them; 55(1)(c) makes them report serious incidents to the AI Office. No provision attaches a number to any of it.
5. So this is a template, not a filing - and that cuts both ways. A vendor's self-measured index is not evidence in your own file. The method is still free to copy.
My Monday: take one agent class in your estate and compute Anthropic's three - coverage, the share of actions a monitor actually sees; review latency; escalation rate. Coverage is the one that will not come back at 100%.
A lab grading itself is still a lab grading itself. This one published the rate at which its own graders disagree.
Which of your agents' actions has nothing ever looked at?
#AIAct #AgenticAI #EUTech #AIGovernance #DevSecOps
AI disclosure: the narration voice, the cover art and the brand ident animation in this video are AI-generated. The script, the claims and the source checks are mine.