Back to the news portal
TechnologyNew analysis today · source 8 October 2026Verified reportAnalysisMulti-source analysisUnited StatesUnited KingdomGlobal

Is AI already building AI?

Anthropic's internal index classified Claude as leading 26% of sampled model-R&D work in August 2026, up from under 1% in February, but classified none of the measured work as fully autonomous. The result is a self-reported measure from one lab—not proof of recursive self-improvement.

By The Impact of AI Editorial DeskReleased 9 October 2026 at 11:15 BST8 min read3 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1Anthropic sampled roughly 15,000 granular R&D tasks and used a six-level scale from no meaningful AI contribution to full autonomy.
  • 2Its index classified Claude as leading 26% of measured model-R&D work in August 2026 and more than 90% as at least collaborative, but classified 0% as fully autonomous.
  • 3The measure covers one company, uses Anthropic models to judge Anthropic work and was frozen against a July task baseline, so it cannot establish industry-wide or sustained recursive self-improvement.
Key themesAI research automationAI agentsMeasurementAnthropicRecursive self-improvementResearch governance

Research topic

Whether internal task-level evidence shows AI systems are taking substantive responsibility for model research, and what measurements would distinguish assistance from autonomous recursive improvement

The Impact of AI analysis cover asking whether AI is already building AI, with a conceptual human-supervised model workshop and incomplete feedback loop; it states 26% AI-led work inside one lab, zero fully autonomous work and that the measure was self-reported.
AI-generated editorial illustration. The researcher, model workshop and incomplete loop are conceptual; they do not depict Anthropic, a real employee, an autonomous laboratory or measured evidence of recursive self-improvement.

The direct answer: AI is doing more R&D work, but people still direct it

Inside Anthropic, Claude was classified as leading 26% of sampled model-research and development work in August 2026, up from less than 1% in February. More than 90% of measured work was rated at least collaborative. None of the sampled task categories reached the study's highest level, where an AI system would execute autonomously without human oversight.

That is material evidence of research automation: the company is reporting that its models now take substantial responsibility for a share of the tasks used to improve future systems. It is not evidence of an AI independently choosing research goals, allocating resources, running an entire programme or repeatedly improving itself without people. Researchers selected the work, supplied context, reviewed outputs and remained accountable for the result.

The State of AI Report 2026 therefore makes a useful distinction. AI is already helping build better AI, but sustained autonomous recursive self-improvement remains unproven. The numbers describe one laboratory's workflow under its own measurement system, not a universal threshold or a forecast of when human supervision disappears.[1][2]

How Anthropic constructed the R&D Automation Index

In July 2026, Anthropic drew a random weekly sample covering 20% of staff in each model-R&D department. Workers decomposed their activity into roughly 15,000 granular tasks. The company organised those descriptions into a hierarchy with 542 nodes and 378 leaf categories, then froze that July baseline so subsequent monthly measurements could be compared against the same map of work.

Each task was assigned one of six levels, from AL0—no meaningful AI contribution—to AL5, autonomous execution without human oversight. Claude agents researched the work descriptions and an independent Claude judge applied the scale. Department results were weighted by sampled person-time, so a common short task did not automatically dominate a rare but lengthy one.

Anthropic compared the automated ratings with blind employee judgments. Exact agreement between the model and humans was 59%, while human-to-human exact agreement was 35%; model and human ratings were within one level 97% of the time. Those figures suggest the scale can capture a broad direction, but they also reveal substantial ambiguity about where assistance ends and leadership begins.[2]

What 26% ‘AI-led’ does—and does not—mean

A task rated ‘led’ can involve an AI agent taking the main operational role while a person frames the objective, checks intermediate work or approves the output. That is different from autonomous research at the programme level. A model might write code, analyse a failed run or propose an experiment without deciding which scientific problem the organisation should pursue or whether the result is safe to deploy.

Aggregation also hides variation. One department may use agents for routine coding while another reserves AI for search or summarisation. The index weights person-time rather than scientific value, so 26% does not mean Claude produced 26% of Anthropic's discoveries or model capability. It cannot be converted into a productivity or progress rate without outcome measures.

The frozen task map improves month-to-month comparison but can become stale as work changes. New AI-assisted practices may not fit July categories, and researchers may alter tasks because agents are available. A rising score could reflect stronger models, better tools, changing management, more favourable task selection or all four.[1][2]

Why self-measurement needs external checks

Anthropic is measuring its own employees with its own models while building the systems whose progress it reports. The company disclosed the design and agreement tests, which is more informative than a vague productivity claim. But there is no independent audit, cross-lab benchmark or public task-level dataset that would allow outsiders to reproduce the 26% figure.

The State of AI Report is also an industry synthesis produced by Air Street Capital, not a systematic review or peer-reviewed study. It selects examples to explain a fast-moving market, and the authors' investment perspective is relevant. Its value here is the cautious synthesis—that automation is real while recursive autonomy is unproven—not independent validation of the underlying company measurement.

Comparisons across organisations would require common definitions. A lab that calls a model ‘leading’ after it drafts most code may be measuring something different from a lab that reserves the label for independently designed experiments. External evaluators also need access to failures, rejected outputs, human correction time and the consequences of mistakes, not only task classifications.[1][2][3]

What this means for researchers, workers and safety teams

For AI researchers, the result suggests that agent literacy and review are becoming core parts of the job. People may spend less time producing first drafts and more time decomposing problems, designing tests and evaluating machine-generated work. That can accelerate iteration, but it can also concentrate attention on the model's preferred approaches and make correlated errors harder to notice.

For other knowledge workers, this is not a direct forecast of automation. Anthropic's staff have unusually strong models, internal tools, compute and support, and their work creates the technology itself. Transfer to medicine, law, education or public administration depends on data access, error costs, regulation and whether outcomes can be checked. A percentage from one AI lab should not be treated as a general employment statistic.

For safety teams, AI-assisted AI research creates a feedback loop even before full autonomy. Faster coding and experimentation can compress the time available for evaluation. Governance therefore needs contemporaneous monitoring: logging agent actions, separating capability work from safety work, requiring human approval for sensitive changes and measuring whether oversight effort keeps pace with automated throughput.[1][2]

What would change the assessment

Confidence would increase if several laboratories applied a shared, preregistered scale to randomly sampled work and allowed independent auditors to inspect anonymised tasks, ratings and corrections. The evaluation should report productivity, quality, failure severity and supervision time alongside automation level. A system that completes more tasks but creates expensive hidden errors is not delivering the same benefit as one that produces reliable work.

Evidence for recursive self-improvement would require a much higher bar: a model would need to originate and execute a sustained sequence of research improvements, demonstrate that the resulting system materially improves the next cycle, and do so with tightly bounded human intervention. Evaluators would need to rule out hidden human selection, cherry-picked successes and simple access to more compute or data.

The State of AI Report published this synthesis on 8 October 2026. On the evidence available, the balanced assessment is that AI agents have become consequential participants in one frontier lab's R&D process. The 0% autonomous result is equally important: it keeps the claim anchored to supervised work and leaves the strongest version of self-improving AI unproven.[1][2][3]

What this means for people

  • AI researchers may shift from producing every artifact to specifying, testing and governing agent-generated work.
  • Workers outside frontier labs should not treat one company's internal task share as a prediction of their occupation's automation.
  • The public has an interest in whether faster model development is matched by independent safety checks and accountable human decisions.

Global context

Frontier AI development is concentrated in a small number of well-resourced United States and Chinese organisations, while its effects are global. One US lab's internal workflow cannot represent research capacity, labour conditions or governance elsewhere. Shared measurement standards and cross-border safety cooperation would make automation claims more comparable without assuming that every institution has the same models, compute or regulatory obligations.

What the evidence does not yet show

  • The 26% estimate comes from one company's internal workflow and has not been independently audited.
  • Anthropic models classified work performed with Anthropic models, creating a potential measurement conflict and shared failure modes.
  • Exact model-human agreement was 59%, showing that automation levels remain judgment calls rather than objective physical measurements.
  • The index measures task responsibility weighted by person-time, not scientific importance, productivity, model improvement or safety.
  • A July task map may miss new work patterns as agent use changes the organisation.
  • The State of AI Report is an investor-produced industry synthesis, not peer-reviewed research.

What to watch next

  • Independent audits and comparable automation indices across multiple AI laboratories.
  • Outcome measures for research quality, failure severity, supervision time and productivity.
  • Whether any task category reaches full autonomy without hidden human selection or correction.
  • Evidence that AI-generated improvements causally accelerate later model generations.
  • Safety evaluation and oversight capacity growing at the same pace as automated experimentation.

Living evidence record

Impact record IAI-11JQVMG

Explore the full tracker

Evidence stage

Observed

Confidence

Supported

Reporting basis

Multi-source analysis

Independent or research support

Not yet

Record status

Monitoring

Last checked

9 October 2026

Source trail

3 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Evidence trail

Sources used for this report

Links checked 9 October 2026

This report is labelled multi-source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Technology

Can language models work without word tokens?

Yes, at research scale. A peer-reviewed Nature study converted existing 1B–8B language models to operate on UTF-8 bytes using 49.1 billion continued-training tokens. The resulting models improved character-level tasks and approached their source models elsewhere, but the evidence is benchmark-based and does not yet prove lower deployment cost or better real-world products.

8 min · 2 sources

Technology

Why do multi-agent AI systems keep failing?

An unreviewed analysis of 22,848 closed issues across 21 prominent open-source projects identified 944 genuine multi-agent problems. Handoffs, execution control and memory dominated—but repository reports cannot establish failure rates in production systems.

7 min · 3 sources

Technology

Does self-hosting keep an AI coding agent away from sensitive source code?

IBM has made a customer-managed version of its Bob development agent generally available, including supported air-gapped deployments. Local inference can keep code inside an approved environment, but the announcement does not prove that every integration, model or generated change is secure.

6 min · 2 sources

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.