ProBackend
ai ai model guardrails
just now7 min read

Anthropic's Three-Tier Cyber Access: A Blueprint for AI Cloud Infrastructure Companies in India

Anthropic folded Project Glasswing and its Cyber Verification Program into a single tiered access framework for Claude's most capable models. The design is a useful study in access governance for any company building on frontier AI.

The real problem Anthropic is solving

Cybersecurity is dual-use. That single sentence explains why Anthropic just rebuilt how defenders get access to its strongest models. The tooling that lets a security engineer find and patch a vulnerability is the same tooling a malicious actor would use to exploit one. So the company made a deliberate choice: its generally available models, including Claude Opus 5.5, Claude Sonnet 5.5, and Claude Fable 5.1, ship with conservative cyber safeguards that block most offensive-style work. Good for the world at large. Rough for the people whose actual job is the offensive part, defensively.

On October 6, 2026, Anthropic announced an expanded Cyber Verification Program that merges two separate experiments into one. For the preceding six months, defenders had reached capable models through two doors. Project Glasswing handed Claude Mythos to organizations securing the most critical software. The original Cyber Verification Program gave vetted teams reduced safeguards on Claude Opus and Sonnet. Now there is one program with three tiers, and every tier routes to the same pool of frontier models: Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1, and whatever ships next.

It's worth reading this less as a product launch and more as an access-governance pattern. The way Anthropic sorted who gets what, and on what proof, is the part I'd want any organization standing on frontier AI to study closely.

Defense Access: the widest door

The bottom tier exists for everyday defensive work. Security operations and incident-response tasks. Reverse-engineering malware. Analyzing and validating vulnerabilities. The eligibility list is genuinely broad: security teams at companies, nonprofits, universities, and government bodies defending systems they own; critical-infrastructure operators of any size, including a regional hospital or a municipal utility; smaller security firms; open-source maintainers; and individual researchers who can point to a track record of reported vulnerabilities.

Because Anthropic expects many defensive teams to qualify here, applications are meant to clear in a few days. That speed matters. The friction an access program adds is a tax on the people you are trying to help, and Defense Access keeps that tax low on purpose.

Red Team Access: authorized offense, hard blocks

The middle tier layers authorized penetration testing and red-teaming on top of everything in Defense Access. Think in-house red teams, government red teams, and security and penetration-testing firms. The scope rule is the load-bearing piece: organizations can only run adversarial testing against systems they are authorized to test, including IT systems in critical industries.

Even here, the guardrails do not go fully away. Anthropic keeps real-time blocks on actions that could cause physical harm or mass disruption, specifically things like deploying ransomware, damaging physical systems, or pen-testing high-risk safety systems. Review runs a few weeks because the bar is higher, and applicants get enrolled in Defense Access while the Red Team application is still in flight. One constraint stands out: this tier is for organizations only. Individual researchers cannot apply.

Specialized Access: the gated room

At the top sits the tier with the fewest cyber blocks, and it is deliberately tiny. It is reserved for a small set of verified organizations authorized to test safety systems that could affect people's lives or disrupt markets. Flight operating systems. Power grids. Telecom networks. Interbank transfer infrastructure. Government administrative networks. The list reads like a catalog of systems where a mistake does not stay digital.

Anthropic says it currently reviews every organization at this tier in depth, in collaboration with the US government. The handover from the old world is clean: existing Project Glasswing members move straight into Specialized Access and do not need reapproval for models they already use. This is what a careful migration looks like. Nobody's running defensive work gets yanked out from under them while the structure changes underneath.

Did the tiers actually work? The CyScenarioBench test

This is the section I'd have wanted before I believed any of it. Anthropic ran Claude Opus 5.5 through CyScenarioBench, an evaluation that measures whether a model can plan and execute multi-stage cyber operations under realistic constraints, with safeguards tuned to each tier. Across five attempts at each of the ten challenges, the pattern was sharp.

With no CVP access, every task was blocked on the very first prompt. In Defense Access, 46 of 50 trials hit a block at some point, and only four completed. In Red Team Access, there were zero blocks, and the model finished 34 of 50 tasks, which lines up with its 67.6% success rate on the same evaluation running with no safeguards at all. Anthropic treats that result as representative of Specialized Access too.

Read those numbers twice. The tiers do exactly what they claim on the tin: they gate, and the higher you climb, the less you are gated. A program that blocks the same work for a hospital's SOC team and a red team testing a power grid is a program that failed the sorting problem. This one appears to pass it.

The Glasswing numbers behind the move

The consolidation isn't happening in a vacuum. Through Project Glasswing, launched in April 2026 with partners including AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA, and Palo Alto Networks, defenders using Claude Mythos reported finding at least 129,000 verified software vulnerabilities between April and July 2026. Anthropic's own open-source scanning turned up another 5,500 between April and October. More than 33,000 of the total were rated critical or high severity.

One honest caveat, which Anthropic states plainly rather than burying: this is a floor, not a ceiling. The figures lean on survey data from only 33 partner reports, and fewer than half of those partners had finished disclosing patched numbers. The company expects the true impact to be at least five times higher. Several partners said the models had compressed months or even years of vulnerability hunting into a shorter window.

The trust question: data retention and EFS

Here is the part a security buyer will ask about before anything else. Enrollment in the program requires data retention, because Anthropic needs to monitor for cyber misuse. You are, by design, trading some data-control for access.

The exit ramp is coming, sort of. A new offering called Enterprise Frontier Safeguards, which combines zero-data-retention privacy with the safeguards, is slated for later in the fall. Once it ships, eligible organizations will be able to store data in cloud infrastructure they control. Until then, organizations that already have Claude Fable 5.1 or Claude Mythos 5.1 on zero data retention can run CVP under zero data retention as well. Access is still narrower on delivery channels: CVP runs on the Claude Platform, Google Cloud's Vertex AI, and Microsoft Foundry, but on Amazon Bedrock only for customers eligible for Enterprise Frontier Safeguards.

A blueprint for AI cloud infrastructure companies in India

Step back from the specifics and you have a template worth stealing. If you're one of the many AI cloud infrastructure companies in India standing up governed access to capable models, Anthropic's tiering answers the question most platform teams fumble: how do you widen access without throwing the safety door open?

The answer it landed on has three legs. Sort by scope of work, not by company size or logo. Tie review effort to actual risk, so a defensive SOC clears in days and a power-grid tester takes weeks under joint government review. And keep real-time blocks on the genuinely dangerous actions even at the top tier, because "trusted" should never mean "unwatched." The friction is graduated, the dangerous edges stay cuffed, and the migration respected people who were already doing the work.

None of this is free. Data retention, organizational-only tiers, and delivery channels that exclude a major cloud are all tradeoffs that some teams will find disqualifying. That tension between openness and control is the whole game, and it's why security researchers have been publicly at odds with vendor guardrails that feel too blunt. A tiered model is one of the more honest attempts to split that difference. Watch how the classifier refinements Anthropic promised actually land over the next few months. That's where programs like this quietly earn or lose defenders' trust.

the real problem anthropic is solving

More blogs