Back to the news portal
Security & DefencePrimary sourceNewsSource analysisUnited StatesNorth AmericaGlobal

Can reduced AI safeguards give cyber defenders an advantage?

Anthropic has opened three levels of less-restricted cyber access and reports 129,000 partner-identified vulnerabilities. Its 50-trial evaluation shows the access controls behave differently, but the effectiveness totals are company-reported and do not establish how many unique flaws were fixed.

By The Impact of AI Editorial DeskReleased 7 October 2026 at 09:57 BST8 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1Anthropic has combined Project Glasswing and its previous Cyber Verification Program into Defense, Red Team and Specialized access tiers, with progressively fewer blocks and stricter eligibility requirements.
  • 2In 50 trials for each setting, the Defense tier blocked 46 trials at some point, while the Red Team tier completed 34 of 50 without safeguard blocks; this tests access controls, not real-world security outcomes.
  • 3The headline 129,000-vulnerability figure is based on partial reports from 33 partners using different triage practices. Fewer than half disclosed patch totals, so discovery volume cannot be read as risk removed.
Key themesCybersecurityAI safeguardsVulnerability researchCritical infrastructureDual-use modelsAccess governance

Research topic

Whether tiered access to less-restricted AI cyber capabilities can improve defensive work without creating an unacceptable route to offensive misuse

The Impact of AI news cover asking whether reduced AI safeguards can give cyber defenders an advantage, illustrated with a shield, code window and controlled access gate.
AI-generated editorial illustration. The shield, code window, access gate and verification cards are conceptual; they do not depict an actual vulnerability, customer system, provider interface or successful attack.

The answer is plausible for verified teams, but the public evidence is not yet independent

Anthropic has expanded access to cyber-capable versions of its models through a three-tier Cyber Verification Program. The practical case is straightforward: a safeguard designed to stop malicious exploitation can also interrupt an incident responder analysing malware or a maintainer validating a flaw in software they are authorised to test. The new system attempts to separate those users through identity checks, organisational verification, security requirements and activity monitoring rather than applying the same restrictions to everyone.

The release provides useful denominators, but it remains a company account of a company programme. Anthropic reports both a controlled evaluation of how often its safeguards block work and large totals for vulnerabilities identified by participating organisations. Neither dataset is presented as a peer-reviewed comparison with other tools or unaided teams. The evidence therefore supports a narrower conclusion: tiered access can materially change what the model is able to do, while the extent of the resulting security benefit and misuse risk still needs independent measurement.[1][2]

Three tiers trade wider defensive access for progressively stronger vetting

Defense Access covers security operations, incident response, malware reverse-engineering and vulnerability analysis. Anthropic says company, university, nonprofit and government security teams, critical-infrastructure operators, smaller security firms, open-source maintainers and individual researchers with a disclosure record may qualify. Red Team Access additionally permits authorised penetration testing. It is limited to organisations, and the model should still block actions capable of physical harm or mass disruption, including ransomware deployment and adversarial testing of high-risk safety systems.

Specialized Access removes the most cyber blocks and is reserved for a limited set of organisations authorised to test systems whose failure could affect lives or markets, including flight operations, power grids, telecommunications, interbank transfers and government networks. Anthropic says it reviews every such organisation in depth with the US government. That is a consequential governance choice: it may provide expert scrutiny, but the public materials do not specify decision criteria, appeal routes, international representation or how consistent judgments will be across jurisdictions.[1][2]

The 50-trial test shows the controls separate access modes

Anthropic tested Claude Opus 5.5 on CyScenarioBench, which contains ten interactive challenges intended to represent multi-stage cyber operations. It ran five attempts per challenge for each access setting, giving 50 trials per setting. Without programme access, all 50 trials were stopped on the first prompt. Under Defense Access, 46 of 50 were blocked somewhere in the challenge and four succeeded. Under Red Team Access, no safeguard blocks occurred and the model completed 34 of 50 tasks.

The company compares that 34-of-50 result with a stated 67.6% completion rate when safeguards are absent, treating Specialized Access as representative of that condition. The comparison suggests the Red Team classifier did not materially suppress benchmark completion. It does not show that 34 operations improved a real organisation's security, nor does it measure false authorisations, analyst judgment, time saved, exploit reliability or downstream harm. Five attempts on ten scenarios are informative about the gate, but too small and controlled to establish a general safety-performance balance.[1]

The vulnerability total needs its denominator and patch status

For Project Glasswing, Anthropic says partners uncovered at least 129,000 verified software vulnerabilities between April and July 2026. Its own open-source scanning found another 5,500 between April and October. More than 33,000 were rated critical or high severity. Those figures are consequential if they represent distinct, valid findings that reached maintainers, because systematic discovery could help reduce risk across widely reused software rather than only inside well-resourced companies.

The published qualification is equally important. The partner total comes from 33 reports, organisations used different triage methods, and fewer than half of partners disclosed how many vulnerabilities were patched. Anthropic describes the count as a lower bound and estimates total impact could be at least five times higher, but that extrapolation is not accompanied by a public sampling method or record-level dataset. Duplicate reports, severity calibration, false positives, disclosure timelines and fixes are not broken out. Finding a flaw is not the same outcome as safely remediating it.[1]

People gain only when findings become fixes without creating new exposure

For maintainers and security teams, fewer false refusals could shorten the path from suspicious code to a reproducible report and patch. Smaller organisations may gain access to analysis that previously required scarce specialist time. People who rely on hospitals, utilities, financial transfers or public services benefit only if the programme produces validated fixes before attackers can exploit the same weaknesses. The absence of a complete patch denominator means the current announcement cannot quantify that public benefit.

Reduced blocking also increases the importance of credential security, audit trails and fast revocation. The programme requires data retention so Anthropic can monitor misuse, with limited zero-retention arrangements and a future Enterprise Frontier Safeguards option. Participants therefore face a second trade-off involving sensitive incident material and provider visibility. A compromised approved account, poorly scoped authorisation or weak internal handling could turn privileged capability into risk, particularly where testing touches operational technology or personal data.[1][2]

What would change the assessment

Confidence would rise with an independent audit that publishes a deduplicated vulnerability denominator, common validation and severity rules, false-positive rates, patch status, disclosure times and comparisons with teams using other tools or no advanced model. A prospective study should record analyst hours, compute and review costs as well as vulnerabilities found. For critical systems, evaluation should test whether access controls withstand stolen credentials, insider threats and ambiguous authorisation without disclosing exploit details that increase harm.

The safety case also needs longitudinal evidence: how many approved users are narrowed or removed, how often real-time monitoring identifies misuse, whether geographic eligibility is equitable and whether material incidents are disclosed. Until those measures exist, the strongest reading is balanced. Anthropic has shown that its tier settings change benchmark behaviour and has reported a large volume of partner findings. It has not yet shown, in independently reviewable evidence, the net reduction in cyber risk produced by opening those capabilities more widely.[1][2]

What this means for people

  • Open-source maintainers and smaller security teams could analyse flaws faster, but each finding still needs expert triage and a safe patch process.
  • Patients, utility customers, bank users and public-service users benefit only when validated findings lead to fixes before exploitation.
  • Security professionals may need to accept provider monitoring and data-retention requirements to use less-restricted capabilities.

Global context

The programme is offered through Anthropic's own platform, Google Cloud and Microsoft Foundry, with more limited Amazon Bedrock availability. Specialized Access is reviewed with the US government even when affected software and infrastructure are global. That makes transparent international eligibility, privacy protections and consistent authorisation standards central to whether the model can serve defenders beyond a small set of well-connected organisations.

What the evidence does not yet show

  • The programme, evaluation and outcome counts are reported by Anthropic; no independent audit or peer-reviewed assessment is linked.
  • The benchmark used ten challenges with five attempts per setting and measures task completion and blocking, not organisational security outcomes.
  • The 129,000 figure comes from partial reports by 33 partners using different triage methods; record-level deduplication and false-positive data are not public.
  • Fewer than half of partners disclosed patched numbers, so discovery totals cannot be converted into vulnerabilities remediated or people protected.
  • Reduced safeguards are inherently dual use, and the public sources do not quantify misuse attempts or failures of participant vetting.

What to watch next

  • Independent validation of the partner-reported vulnerability and severity totals.
  • A common public denominator for unique findings, accepted disclosures, fixes, duplicates and false positives.
  • Evidence that tier vetting, monitoring and revocation prevent misuse in practice.
  • Whether qualified defenders outside the United States receive comparable access and review treatment.

Living evidence record

Impact record IAI-0098Q27

Explore the full tracker

Evidence stage

Announced

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Not yet

Record status

Monitoring

Last checked

7 October 2026

Source trail

2 direct sources across 2 source types.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

Evidence trail

Sources used for this report

Links checked 7 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Security & Defence

Are cyber budgets rising faster than AI safeguards?

New analysis today of PwC’s 1 October survey of 3,934 business and technology leaders across 71 countries. Spending expectations are rising and attacks on AI systems top respondents’ preparedness concerns, but the findings are self-reported perceptions—not audited resilience or incident rates.

7 min · 2 sources

Security & Defence

Can AI detect attack types it never saw in training?

A hybrid Transformer–LSTM detected two held-out categories in the UNSW-NB15 benchmark, with 88.13% recall on the combined unseen subset. It was a binary, row-level laboratory test—not evidence of live zero-day defence.

8 min · 1 source

Security & Defence

Will Apple’s new consent controls make AI agents safer on Macs?

Apple says it will add controls requiring very explicit user action before an app receives Full Disk Access, warning that autonomous AI raises the risk of exposing files, mail, messages and browsing history. The direction is important, but Apple has not yet published the interface, release version, rollout date or evidence that the design prevents mistaken consent or misuse after permission is granted.

10 min · 5 sources

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.