Anthropic announced on Tuesday a sweeping revamp of its Cyber Verification Program (CVP), folding two initiatives it has run for the past six months — Project Glasswing and the original CVP — into a single three-tier scheme that gives vetted security professionals access to its most capable models with progressively fewer safeguards.
The results so far are substantial. Organizations participating in Project Glasswing, which gives teams securing critical software access to the Claude Mythos family, identified at least 129,000 verified vulnerabilities between April and July. Anthropic's own open-source scanning found another 5,500 between April and October. More than 33,000 of the vulnerabilities discovered so far have been rated critical or high severity. The company cautions that these figures are likely a significant undercount: because the data comes from a limited survey of partners, it estimates the true impact could be at least five times higher.
The new structure has three tiers, all of which include access to Claude Opus 5.5, Sonnet 5.5, Mythos 5.1 and future models. The Defense tier covers incident response, malware analysis and related defensive work, and is open to security teams, critical infrastructure operators, open-source maintainers and researchers with a track record of reported vulnerabilities; Anthropic aims to respond to these applications within days. The Red Team tier adds authorized penetration testing and adversarial operations, is open to organizations only, and reviews are expected to take a few weeks, with applicants receiving Defense-level access in the meantime.
The Specialized tier carries the fewest restrictions and is reserved for a small group of organizations authorized to test safety-critical systems — power grids, flight operating systems, interbank transfer infrastructure. Anthropic vets each member of this tier in coordination with the US government, and existing Glasswing members will transition into it without reapproval for current models.
How much the leashes actually loosen is quantified — by Anthropic itself, in company-run tests that have not been independently verified. On its CyScenarioBench suite of multistage offensive cyber operations, 50 Claude Opus 5.5 trials were run per tier. Without CVP access, all 50 were blocked at the first prompt. Under Defense access, 46 of 50 trials encountered blocks and four succeeded. Under Red Team access, zero trials were blocked and 34 succeeded — which Anthropic says is effectively equivalent to the 67.6% success rate it records with no safeguards at all. The company stresses this measures permission boundaries on synthetic challenges, not real-world misuse rates.
Some guardrails are absolute: real-time blocking still applies to actions that could cause physical harm or mass disruption, including ransomware deployment and testing of high-risk safety systems. Enrollment generally requires data retention so Anthropic can monitor for misuse, though its planned Enterprise Frontier Safeguards — due later this fall — will let eligible organizations keep data inside cloud infrastructure they control, and existing zero-data-retention arrangements on Fable 5.1 or Mythos 5.1 can carry over into the program.
Access is available through the Claude Platform, Google Cloud's Vertex AI and Microsoft Foundry, with Amazon Bedrock limited to customers eligible for Enterprise Frontier Safeguards. The expansion lands against a backdrop of lingering unease: when Claude Mythos Preview launched in April, it stoked fears that AI systems could find and exploit software flaws before defenders had patched them. Anthropic's answer is essentially to co-opt the people most worried about offensive AI — by giving them the same tools first.
Comments (0)
Log in to join the discussion
Log InNo comments yet