The Dirty Secret Every CISO Knows
Dev Rishi walked into a room of about fourteen CISOs at Anthropic's roundtable and asked whether everyone had their AI governance policies written down. Every hand went up. Then he asked who was actually enforcing them. "Everybody chuckled," he told the VB Transform 2026 audience in Menlo Park. "It was like the dirty secret in the room that everyone has these policies, but no way to actually make them real."
That gap between policy and enforcement is exactly why Rubrik built SAGE — the Semantic AI Governance Engine — and why it now sits between every agent action and its execution. Rubrik's GM of AI put the system in production on itself first. Their agents run in YOLO mode, which strips the permission prompt out of workflows entirely. No human clicks approve. Instead, a second AI judges every move in real time against the policies written in natural language.
The bet is simple and uncomfortable: if you ask an agent to act autonomously, it will. The real question is whether you should let it, and who decides when the line gets crossed.
Who Watches the Watcher — And How Well Does It Watch?
Rishi didn't stumble into this role by accident. He ran Predibase, the generative AI infrastructure startup Rubrik acquired in June 2025, before that led the ML product team at Google that became Vertex AI, and was Kaggle's first product manager as it grew from roughly one million to ten million users. Harvard computer science, both degrees. He knows how these systems behave because he's been building them for a decade.
So when he walked into 200 customer conversations over his first three and a half months at Rubrik — spanning what looked like the Global 2000 — he wasn't fishing for complaints. He asked about cost, latency, performance, orchestration. "Pretty consistently, what I heard through all of those conversations was that all of those are pretty secondary," he said. "The main challenge is actually, how do I get this approved from a security and risk standpoint? I'm concerned about all the different things that could go wrong."
That concern was, he put it, "one of the biggest things constraining ROI." And the data backs him up. A VentureBeat Pulse survey presented at the same event found that 66% of enterprises already allow or are actively building toward production deployment with zero human review. Yet only 5% fully trust the automated evaluations that would make that decision. Five percent.
Rishi's timing has a market behind it. Eighty-two percent of enterprises still name their primary AI provider's built-in guardrails and cloud controls as their main agent security layer. Fifty-nine percent plan to adopt, add, or replace agent security tooling within the next 12 months. Only 12% include an agent-identity product in what they're considering, even with credential sharing still the norm. Every CISO at that Anthropic roundtable had a policy document and no enforcement mechanism. Rubrik built a product for the space between the two.
Backtesting: The Feature Nobody Asked For (But Probably Needs)
Rubrik Agent Cloud reached general availability in February 2026, though not everything Rishi described ships in it yet. Backtesting is just starting to roll out, and it's the feature that made him pause mid-sentence during the fireside chat.
Backtesting replays an organization's historical agent actions and tool calls against a new policy, showing where the policy would have stepped in and where an action would have sailed through uncaught. You edit the policy in real time. He called that archive "one of the most valuable data troves an enterprise holds."
Real-time detection and blocking turn out to be the entry point rather than the whole product. Some attacks never trip a single-action rule. "No individual turn of the conversation was problematic, but if you took the session as a full trace, that ended up being problematic," Rishi said.
Agent Cloud runs batch analysis across entire session traces every hour or every day and surfaces what Rubrik calls insights — the problems no individual guardrail caught. The same Zero Labs report found that 88% of enterprises say they lack the ability to roll back agent actions without system disruption. That recovery gap sits squarely in Rubrik's original line of business.
The attacks no single turn reveals are the ones keeping security teams up at night. A financial services company Rishi met the morning of the session made the point for him. None of the individual permissions an agent has are bad on their own, he said. The agent needs every one of them to do its job. "It should have permission to each of those systems, but it's the combination that ends up becoming really destructive." Traditional identity and access management never priced in this combination because it relied on the judgment of the employee holding the credentials. Agents supply none.
The Accuracy Question Nobody Answered
A skeptical CISO will ask the question the fireside did not: what is SAGE's false positive rate? Its false negative rate? Rishi offered no metrics for the judge itself.
SAGE is a non-deterministic model policing other non-deterministic models, and the closest thing the architecture gives to an answer is auditability. Backtesting and the batch insights both leave a human-reviewable trail of each call SAGE made and whatever got past it. Who watches the watcher, for now, is a trail of receipts rather than a benchmark.
Until that benchmark exists, AI in the loop stays an operational wager rather than a quantified control. The architecture bet is that models good at understanding language can police other models. It's a bet Rishi made with "a lot of naivety and innocence, honestly." His team of AI infrastructure people took on a problem that security engineers own, and SAGE became the answer.
Three questions fall out of the session for security teams. How many of the guardrails now in production depend on a human clicking approve, and what happens to that workload as agent count grows? Does anything in the stack enforce semantic intent, or is it all allow and deny lists? And can the team backtest agent behavior against a new policy, then unwind a multi-turn session without taking systems down?
The Lethal Trifecta
Simon Willison coined the term in June 2025, and Rishi pointed to it as the attack pattern that keeps him up at night. The lethal trifecta describes an agent that holds private data while taking in content nobody vetted, with a channel to send what it finds to the outside world.
The danger isn't any single permission. It's what happens when individually legitimate permissions stack. An agent granted Salesforce access and email access on an employee's credentials has done nothing wrong yet — with yet being the operative word. "A very simple example is that an agent can start pulling data from Salesforce and then decide to accidentally leak and exfiltrate that out via an email," Rishi told the audience.
A financial services company made the point for him the morning of the session. None of the individual permissions an agent has are bad on their own, Rishi said. The agent needs every one of them to do its job. "It should have permission to each of those systems, but it's the combination that ends up becoming really destructive."
Traditional identity and access management never priced in this combination because it relied on the judgment of the employee holding the credentials. Agents supply none.
The pattern shows up in the data too. A separate VentureBeat June Pulse survey of 107 qualified enterprise respondents found that 69% run credential sharing somewhere in their agent fleet. Companies with shared credentials anywhere reported a security incident or near-miss at a 63.5% rate (47 of 74), versus 40.9% (9 of 22) where every agent carries its own scoped identity.
Cutting agents off from public resources entirely would defeat their purpose. The problem returns to adjudicating intent in context rather than revoking access.