Lead
Anthropic has rolled out a tiered access system that lets vetted cybersecurity teams run Claude with far fewer guardrails — a move timed to new numbers from the company's Project Glasswing vulnerability-hunting initiative, which says it verified more than 129,000 software vulnerabilities between April and July 2026 alone, with another 5,500 confirmed through October via open-source scanning.
What happened
The new framework, called the Cyber Verification Program (CVP), sorts approved organizations into three tiers:
- Defense Access — incident response, malware reverse engineering, vulnerability analysis and validation.
- Red Team Access — adds authorized penetration testing and red-teaming for defensive purposes.
- Specialized Access — the fewest safeguards, limited to a small set of organizations vetted to test safety systems directly.
Each tier covers Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1, and future models. Anthropic's internal benchmark (CyScenarioBench) illustrates the gap: on Claude Opus 5.5, standard safeguards blocked 46 of 50 adversarial tasks under Defense Access, while Red Team Access let the same model complete 34 of 50 — the same completion rate as running with no safeguards at all. Outside the program, every task was refused at the first prompt.
Of the vulnerabilities Project Glasswing has verified, Anthropic says over 33,000 are rated critical or high severity — and calls that a likely undercount, since the figure is based on reporting from only a subset of participating partners. The company believes the real total could be at least five times higher.
Why it matters
Independent analysis adds useful context on actual risk. VulnCheck researcher Patrick Garrity found that of 300 Project Glasswing-attributed vulnerabilities reviewed, only two — 0.67% — have seen confirmed exploitation in the wild: an SQL injection in Ghost CMS (CVE-2026-26980) and a session-forgery flaw in Rejetto HTTP File Server (CVE-2026-61500). The rest span 39 critical, 141 high, 81 medium and 18 low-severity findings with no observed active exploitation so far.
That gap is the real headline: AI is clearly lowering the cost of finding vulnerabilities at scale, but discovery volume doesn't equal exploitability or real-world impact. The same dual-use tension cuts the other way too — separate testing from 1Password and Veracode found that AI-generated patches can introduce their own fresh vulnerabilities. Veracode reports that roughly 44% of AI code-generation tasks introduced a risky security flaw in its tests, with overall security pass rates holding flat around 56% (versus 55% previously) even as the volume of AI-generated code in production pipelines keeps climbing.
What to do
- Treat AI-assisted vulnerability discovery as a lead-generation signal, not a severity verdict — triage against exploitability, exposure, and reachability before reprioritizing patch cycles.
- If your organization adopts AI coding assistants, pair them with mandatory SAST/SCA gating in CI — a flat ~56% pass rate on AI-generated code means unreviewed merges are roughly a coin flip on introducing new flaws.
- Check exposure to the two confirmed CVEs if you run Ghost CMS or Rejetto HTTP File Server, and patch immediately if you haven't.
- If you're requesting elevated AI red-team access anywhere, make sure your own vetting, logging and oversight controls match the tier of capability you're being granted — reduced safeguards shift risk onto your governance, not away from it.
