THE SIGNAL IN ONE SENTENCE
Anthropic has published three kinds of measurements from inside the process it uses to build AI: how much research and development Claude performs, how the company monitors research agents, and how it divides computing power between capability and safety work. Its prototype index says Claude led 26 percent of measured AI R&D in August, meaning the model completed most of those tasks from a high-level prompt while a human supervised. More than 90 percent reached at least the lower level called collaboration. Claude was not fully autonomous in any measured category. Anthropic also says about 30,000 agents were active at any one time on its most-used internal research platform, every action passed through an online monitor before execution, and roughly one in 47,000 of more than a billion August decisions was blocked. In a separate one-week July snapshot, the company classified about 6 percent of AI R&D compute as safety work and about 12 percent of AI-driven AI R&D compute as safety work. Those numbers are useful because they give outsiders something sharper than the phrase “AI is helping build AI.” They are not an independent audit, a failure rate, or proof that Claude is autonomously improving itself. Anthropic built the index, supplied the internal evidence, used Claude to classify much of the work, and published the result. The plain signal is that the lab has put gauges on a very fast machine. The next job is to calibrate those gauges outside the company, publish the denominators, and connect more automated research to measured capability gain and real safety outcomes.
01
WHAT ACTUALLY CHANGED
Anthropic published a prototype R&D Automation Index built from a map of model-development work inside the company. Using Epoch AI's six-level scale, it defines “collaborates” as completing large pieces under close human direction and “leads” as completing most of a task end to end from a high-level prompt while a human supervises. As of August, the company says Claude led 26 percent of measured work, more than 90 percent reached at least collaboration, and no measured subset was fully autonomous.
The reported trajectory is steep. Anthropic's chart places the share led by Claude below 1 percent in February and at 26 percent in August. The underlying basket came from about 15,000 granular tasks extracted from sampled staff work records, organized into a fixed tree with 378 leaf categories, then weighted using person-time as a rough proxy for importance. That is a serious methodology and still a prototype built from company records by company models.
Anthropic also disclosed the scale of its internal agent system. It says approximately 30,000 agents were doing research and engineering work at any one time on the company's most-used internal platform in August. Every action on that platform passed through an online monitor before execution and was also ingested for offline review. The company reports that the online monitor blocked 0.002 percent of more than a billion decisions, or about one in 47,000.
The oversight disclosure includes a human-review funnel. Anthropic says offline monitoring flags roughly 100,000 transcripts each week for additional classifier review, with about 50 highest-priority cases reaching people. Humans review online-blocked actions within one week. These counts describe the monitor's workflow. They do not establish how many harmful actions the monitor missed, how many blocks were false alarms, or how severe the stopped actions would have been.
A third measurement follows compute. For July 13 through July 20, Anthropic says about 6 percent of compute assigned to AI R&D went to safety and about 12 percent of compute assigned to AI-driven AI R&D went to safety. The company calls those estimates conservative and excludes mixed work that advanced safety and capability equally. It also says compute is an imperfect proxy because important safety research can require more human thought than accelerator time.
02
WHY THIS MATTERS
Automation inside an AI laboratory can compress the feedback loop. A model can help write code, run experiments, diagnose defects, analyze results, and prepare the next experiment. If the work is good, a more capable model may arrive sooner and become a stronger research assistant for the following cycle. That does not require a science-fiction system that secretly rewrites itself. A supervised production loop can accelerate long before it becomes autonomous.
The label on the task changes the story. Leading 26 percent of weighted work does not mean Claude made 26 percent of research decisions, produced 26 percent of capability gains, consumed 26 percent of compute, or removed people from 26 percent of jobs. One model-led task can be routine and another decisive. The index needs to remain attached to its task map, weighting rules, evidence, model version, human interventions, and uncertainty.
Monitoring coverage is not detection coverage. Sending every action through a checker is valuable, just as scanning every bag at an airport is valuable. It does not prove the checker recognizes every prohibited item. A rate of one blocked action in 47,000 is therefore neither a safety score nor a misconduct rate. It is the output of one monitor under one policy on one platform, with unknown false negatives and a denominator containing many ordinary actions.
Compute allocation is a governance clue, not a virtue meter. Six percent can be compared over time only if the definitions, workload boundaries, accounting system, and treatment of mixed projects remain stable. It cannot tell readers whether the safety work was good, whether a low-compute insight changed a release, or whether safeguards deployed elsewhere were omitted. The allocation matters most when paired with decisions and outcomes.
Public measurement can become infrastructure. A regular, versioned series from several laboratories, checked by independent evaluators using compatible definitions, would let governments and researchers see whether oversight and safety work keep pace with automated development. It could also expose inconvenient changes that a single launch-day safety report can smooth away. Anthropic has proposed that possibility. Comparable evidence is the test.
03
WHERE IT COULD HELP
- Publish the complete measurement package on a fixed schedule: task taxonomy, weights, sampling method, model and judge versions, human-rating protocol, disagreement intervals, intervention rules, exclusions, revisions, and historical values that never disappear when the dashboard changes
- Give independent evaluators secure access to a representative sample of tasks, source evidence, agent traces, monitor decisions, blocks, human reviews, false positives, known misses, and capability evaluations, with authority to publish disagreements and methods without company approval
- Separate five quantities that are easy to blur: the share of tasks where AI assists, collaborates, or leads; the share of consequential decisions humans retain; the compute used; the measured quality and speed of outputs; and the capability gain attributable to the automated loop
- Stress-test the monitors with planted prohibited actions, novel attack paths, cross-agent coordination, slow-burn misconduct, corrupted context, compromised tools, and deceptive but plausible rationales, then publish detection, false-alarm, latency, severity, and human-response results
- Create a cross-laboratory reporting standard with common definitions, signed data snapshots, third-party calibration, change logs, incident denominators, and release triggers so a rise in automated AI R&D can be compared across time and companies rather than interpreted from marketing language
KEEP A HAND ON THE WHEEL
These are Anthropic's measurements of Anthropic's work. They have methodological detail and important limits. Claude helped construct the task tree, research automation levels, and judge the evidence. Anthropic says the judge exactly matched human ratings 59 percent of the time while humans exactly matched one another 35 percent of the time, and that model and human ratings were within one level 97 percent of the time. Borderline classifications remain contestable. The index uses sampled July work records and person-time as a rough importance weight, then freezes a task tree while practices keep changing. Thirty thousand concurrently active agents are not thirty thousand independent researchers or autonomous successor builders. The one-in-47,000 block rate measures what a particular online monitor stopped, not all harmful behavior, failures, or undetected actions. Full monitoring coverage does not prove full detection. The compute figures cover one week, rely on internal categories, omit some safeguards compute, and do not measure research quality or safety outcomes. Anthropic says Claude was not fully autonomous in any measured AI R&D subset. Watch for repeated time series, an external audit, known-misbehavior detection rates, false-negative tests, cross-lab replication, stable compute definitions, evidence connecting automation to capability gain, and proof that human supervision remains practical as the loop speeds up.
04
TERMS WORTH KEEPING
OPEN GLOSSARY CARD
AI-led R&D
Research and development work in which an AI system completes most of a defined task from a high-level instruction while a person supervises the result.
OPEN GLOSSARY CARD
Automation level
A category describing how much of a task an AI performs and how much direction, intervention, or completion work remains with people.
OPEN GLOSSARY CARD
Detection coverage
The share of relevant harmful or prohibited behavior that a monitoring system can correctly recognize, including cases deliberately planted to test it.
SOURCES AND VERIFICATION STATUS
This article was written from the materials below. Product claims and dates were checked against those sources on September 18, 2026.
PUBLICATION RECEIPT: Revision 1. Published September 18, 2026.
THE PUBLICATION ENGINE
WANT A SIGNAL OF YOUR OWN?
We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.
WORK WITH US