Anthropic told the public something most AI labs have kept behind closed doors: how much of its own research is now done by its own AI. On September 17, 2026, the company said Claude “leads” 26% of the artificial intelligence research and development work happening inside Anthropic, and that the model contributes at some meaningful level to more than 90% of that work. The disclosure came through a new internal measure Anthropic calls the R&D Automation Index, and the company said it plans to publish updated figures on a regular schedule.

The announcement, first reported by Reuters and picked up within days by outlets including Unite.AI, The Straits Times, Newser and the Economic Times, lands at a moment when frontier labs are under pressure to show their work rather than just their benchmarks. For an industry that has spent two years arguing about how close AI is to improving itself, a named company putting a number on it is a genuine shift in tone. This is a Claude AI R&D story, but it is also a preview of a metric every major lab will eventually be asked to publish.

What Anthropic Actually Disclosed on September 17

Strip away the framing and the core claim is narrow and specific: as of August 2026, Claude “leads” 26% of Anthropic’s model research and development work. Anthropic defines “leads” as being able to complete most of a given task end-to-end from a high-level prompt, while a human still supervises the outcome. That is a meaningfully higher bar than a coding assistant autocompleting a function. It means a researcher can hand Claude a research goal and get back a mostly finished result, not just a draft.

Anthropic paired that headline figure with a second, larger one: work at the “collaborates” tier or above, meaning Claude handles large chunks of a task under close human direction, covers more than 90% of R&D activity at the company. Reporting tied to the same disclosure also cited roughly 30,000 AI agents active on Anthropic’s internal platform during August, a rough proxy for how much of the day-to-day engineering grind now runs through automated systems rather than a person opening a terminal. None of this, Anthropic was careful to say, adds up to full autonomy. The company stated plainly that Claude is not operating fully autonomously in any measured subset of AI R&D work.

Inside the R&D Automation Index: Leads vs. Collaborates

The R&D Automation Index is the actual product here, more than the 26% number itself. It is Anthropic’s attempt to turn a vague, often-hyped question, how much is AI doing to build the next AI, into something with a repeatable definition and a publication cadence. The first results came out of an Anthropic Institute post dated September 17, 2026, and Anthropic described them as a prototype, meaning the methodology could change as the company refines what counts as “leading” a task versus merely “assisting” with one.

Automation TierAnthropic’s DefinitionShare of R&D Work (Aug. 2026)
LeadsCompletes most of a task end-to-end from a high-level prompt, under human supervision26%
Collaborates or higherHandles large chunks of work under close human directionMore than 90%
Fully autonomousNo human supervision required for the task0% (not observed in any measured area)

Read the two rows together and the picture is less dramatic than “AI builds AI” headlines suggest, and more useful. Claude is doing a lot of the heavy lifting, but a human is still the one deciding whether the output ships. That gap between “leads” and “fully autonomous” is exactly the distinction Anthropic wants regulators and researchers to notice.

What Claude Actually Does Inside Anthropic’s Research Pipeline

Anthropic’s own framing describes “AI research and development work” broadly rather than listing individual job categories, but the “end-to-end from a high-level prompt” language points to the kind of work that used to fill a research engineer’s week: standing up an experiment, writing and iterating on code, running an evaluation suite, and summarizing what happened. The roughly 30,000 concurrent AI agents figure tied to the same disclosure suggests this isn’t a handful of pilot projects. It is closer to a standing internal workforce of automated processes running alongside human researchers, most of it supervised rather than left to run unattended.

For readers tracking Anthropic’s broader model lineup, this operational shift sits alongside the company’s other 2026 moves, including its Claude Fable 5.1 release and its Mythos platform changes. A company that is using its own models this heavily inside its research process has a direct incentive to keep improving the tools its own engineers depend on daily.

Timeline: How the Disclosure Spread

Reuters broke the story on September 17, 2026, under the headline describing Claude as leading “a quarter of work building its next AI models.” The wire report was mirrored the same day by Yahoo Finance. Within 24 to 48 hours, secondary coverage followed a familiar pattern for a fast-moving AI story: Unite.AI and The Straits Times ran explainer pieces on September 17 and 18 that dug into the R&D Automation Index methodology, framing it against long-running “AI systems moving toward building themselves” concerns. By September 18 and 19, outlets including Newser, Dataconomy and the Economic Times’ CIO desk had folded the figures into broader pieces about frontier labs and self-improvement.

That spread, wire service to AI trade press to enterprise-tech press within three days, is itself a data point. It shows how quickly a single self-reported metric from one lab becomes an industry reference number, whether or not the methodology behind it has been independently audited.

The Recursive Self-Improvement Question

The reason a 26% figure generated this much coverage has little to do with the number itself and everything to do with what it implies. AI safety researchers have spent years debating recursive self-improvement, the scenario where a model gets good enough to meaningfully accelerate work on its own successor, creating a feedback loop that speeds up with each generation. Anthropic’s own coverage explicitly ties the R&D Automation Index to that debate, describing it as a way to show how close frontier labs are to recursive self-improvement rather than just a productivity metric.

Anthropic’s answer, based on what it disclosed, is: closer than a lot of people assumed, but not there yet. Zero percent of measured R&D work runs without human supervision. That is the load-bearing sentence in the whole disclosure. It lets Anthropic claim transparency and momentum at the same time it draws a hard line against the scarier version of the story. Whether that line holds as the “leads” percentage climbs in future index updates is the open question the company has now committed to answering in public, on a schedule.

How Anthropic Compares to OpenAI and Google DeepMind

Here is where the story gets genuinely interesting for anyone tracking the competitive landscape: neither OpenAI nor Google DeepMind has published a comparable quantitative figure for how much of their own R&D is AI-led. Coverage of the Anthropic disclosure repeatedly frames it in the context of frontier labs broadly, but no outlet reviewed alongside this story cites an equivalent percentage from either rival. That makes Anthropic first to put a number on the table, for better or worse.

LabPublic AI-Led R&D MetricStatus as of Sept. 20, 2026
Anthropic26% “leads,” 90%+ “collaborates or higher”Disclosed Sept. 17, 2026, updated regularly
OpenAINone publishedNo equivalent index identified in current coverage
Google DeepMindNone publishedNo equivalent index identified in current coverage

That gap does not mean OpenAI and DeepMind aren’t using their own models internally at a similar or greater scale. It means they haven’t chosen to quantify and publish it the way Anthropic just did. Shattered.io has separately covered Google DeepMind’s own public statements about the pace of self-improving AI, and OpenAI’s GPT-6 Astra training run, both of which speak to how much compute and internal automation the top three labs are already pouring into their own next-generation systems, even without a published percentage attached.

Historical Context: From Autocomplete to Co-Researcher

It is worth remembering how recently “AI helps build AI” was a novelty rather than a disclosed metric. GitHub Copilot popularized AI-assisted coding in 2021, mostly as line-by-line autocomplete. By 2023 and 2024, coding agents could handle small, well-scoped tickets with light supervision. The jump Anthropic is now describing, a model completing most of a research task end-to-end from a single high-level prompt, represents a different tier of capability that simply did not exist industry-wide three years ago.

Anthropic’s talent pipeline has moved in step with that shift. The company’s recent hire of former Meta AI researcher Andrew Tulloch is one example of frontier labs competing hard for the people who build and tune the very systems now doing a quarter of the lab’s own R&D. The market for researchers who can supervise, not just train, these systems has become its own hiring battleground.

Market and Industry Reaction

Anthropic is privately held, so there was no stock move to track directly off the announcement, and none of the outlets that covered the story reported a market reaction tied to Anthropic itself. The more interesting reaction has been narrative rather than financial: the story got folded almost immediately into the wider 2026 conversation about AI safety oversight, with several outlets noting that AI developers face growing pressure from regulators, companies and researchers to prove their systems are reliable before capability claims like this one get taken at face value.

That pressure matters because Anthropic has also had a rockier few months on the security side. The company’s disclosure of AI-led R&D work follows its own admission of a fourth Claude-related cyber breach in 2026, a track record that gives skeptics an easy rebuttal whenever Anthropic asks the public to trust a self-reported capability metric: the same systems doing more of the company’s R&D are also the systems tied to its most serious security incidents this year.

What the 26% Figure Means for Software Engineers

For working engineers outside Anthropic, the practical takeaway isn’t that AI is about to replace research teams. It’s that “leads” level automation, a model taking a high-level goal and returning a mostly finished result with only supervision-level oversight, is now something a major lab says it is running in production internally, not just demoing. That is a different bar than most teams outside frontier labs are working with today, where AI coding tools mostly still operate at the “collaborates” tier: useful for large chunks of a task, but not trusted to run the whole thing unattended.

Anthropic’s own numbers suggest the gap between those two tiers, 26% fully leading versus 90% collaborating, is where most of the near-term productivity gains actually sit. Engineering teams evaluating AI-assisted workflows should treat “collaborates” as the realistic, achievable target for 2026 and 2027, and treat “leads” as the tier vendors will market aggressively but that still requires the kind of internal tooling, evaluation harnesses and human review Anthropic describes building around Claude before trusting it with an end-to-end task.

The Risks Anthropic Itself Is Flagging

Give Anthropic credit for not soft-pedaling the concern baked into its own announcement. By explicitly tying the R&D Automation Index to recursive self-improvement, the company is naming the exact fear researchers have raised about frontier AI for years: that once a lab’s models are good enough to meaningfully speed up the next model, the pace of progress stops being paced by human researchers at all. Anthropic’s answer for now is oversight, human supervision on every measured task, a hard zero on full autonomy, and a promise to keep publishing the number so outsiders can watch whether that changes.

Whether a self-published, self-defined index is sufficient oversight is a fair question, and one none of the coverage around this story has fully answered. There is no independent auditor validating Anthropic’s tier definitions, and “leads” versus “collaborates” is Anthropic’s own taxonomy, not an externally standardized one.

In Anthropic’s Own Words

Anthropic’s own published framing of the finding is direct: Claude “leads” 26% of Anthropic’s AI R&D work, according to the Anthropic Institute’s post on measuring the pace of AI development. The company positioned that number as the headline finding of the first R&D Automation Index run, covering internal activity through August 2026.

Predictions: Where AI-Led R&D Disclosure Goes From Here

A few things look likely to follow from this disclosure over the next two to three quarters:

  • Anthropic’s next R&D Automation Index update will almost certainly show the “leads” figure climbing above 26%, since the company built the metric specifically to track and publicize that trend.
  • Competitive pressure will push at least one of OpenAI or Google DeepMind to publish its own version of this metric within the next year, even if the methodology differs from Anthropic’s.
  • AI safety researchers and policymakers will push back on self-reported figures and start asking for independent verification of automation-tier definitions rather than taking lab-published numbers at face value.
  • Expect “R&D Automation Index” or similar phrasing to show up in AI safety summit agendas and congressional testimony before the end of 2026, given how directly it touches the recursive self-improvement debate.
  • Enterprise AI vendors will start marketing “leads”-tier automation as a selling point, even though Anthropic’s own data shows that tier still represents roughly a quarter of even a frontier lab’s most sophisticated internal workflows.

Why This Story Outlasts a Single News Cycle

Plenty of AI announcements fade within a week. This one is more likely to stick because Anthropic explicitly committed to a publication cadence, meaning the 26% figure isn’t a one-off flex, it’s a baseline the company will be measured against every time it reports again. That turns a single news story into a recurring one: the next Anthropic Institute update becomes, by design, a comparison point against September 2026’s number, and every rival lab’s silence on the same metric becomes more conspicuous the longer it continues.

Frequently Asked Questions

What did Anthropic actually announce on September 17, 2026?
Anthropic said Claude “leads” 26% of the company’s AI research and development work, and contributes at the “collaborates” level or higher to more than 90% of that work, based on a new internal measure called the R&D Automation Index.

What does “leads” mean in Anthropic’s R&D Automation Index?
Anthropic defines “leads” as Claude completing most of a given task end-to-end from a high-level prompt, while a human still supervises the result. It is not full autonomy.

Is Claude fully autonomous in any part of Anthropic’s research process?
No. Anthropic stated explicitly that Claude is not operating fully autonomously in any measured subset of its AI R&D work as of the August 2026 data behind this disclosure.

Have OpenAI or Google DeepMind published similar figures?
Not as of September 20, 2026. Coverage of Anthropic’s disclosure frames it in the context of frontier labs broadly, but no equivalent quantitative metric from OpenAI or Google DeepMind appears in current reporting.

How many AI agents does Anthropic run internally?
Reporting tied to the same disclosure cited roughly 30,000 AI agents active on Anthropic’s internal platform during August 2026, offered as context for the scale of automated activity behind the R&D figures.

Why does this connect to fears about recursive self-improvement?
Recursive self-improvement describes a scenario where AI models meaningfully speed up work on their own successors, creating a feedback loop. Anthropic’s own coverage ties the R&D Automation Index directly to measuring how close frontier labs are to that scenario.

Will Anthropic keep publishing this metric?
Anthropic said the R&D Automation Index results are a first, prototype run and that it intends to publish updated measures regularly going forward.

Is this metric independently verified?
No independent auditor has validated Anthropic’s tier definitions or figures. The 26% and 90%+ numbers are self-reported by Anthropic through its own Institute publication.