ProBackend
ai agent safety failures
1 hour ago6 min read

Why the NIST AI Risk Framework Is the Wrong Shape for the Misalignment Problem

As AI misalignment reports pile up, enterprises are reaching for the NIST AI RMF. But a closer look reveals a gap between what the framework actually does and what frontier AI safety demands.

The Misalignment Conversation Has Outgrown Its Tools

Something shifted in the last eighteen months. AI safety moved from a fringe concern — the domain of a few researchers getting side-eyed at conference panels — to a board-level checkbox. Reports of models exhibiting misaligned behavior, doing things their operators didn't intend and sometimes didn't expect, have accumulated to the point where you can no longer wave them away as anomalies. Our own analysis of what the Hugging Face breach really means is one of several post-mortems that made the abstract concrete for security teams.

The response from the enterprise world? A frantic scramble for structure. Frameworks. Checklists. Anything that makes the problem look manageable.

And the first thing everyone grabs is NIST's AI Risk Management Framework. Fair enough — it's the only government-backed option with a public website, a playbook, and a roadmap page that doesn't require a PhD to navigate. But here's the thing: the AI RMF wasn't built for this. It wasn't built for misalignment at all. It was built for something more mundane, and understanding that gap explains a lot of the frustration CISOs are feeling right now.

What the AI RMF Actually Is

Let's start with the facts, because the marketing around this framework sometimes obscures them.

The NIST AI Risk Management Framework (AI RMF 1.0) was published on January 26, 2023, authored by Elham Tabassi under the NIST AI 100-1 report designation. It was directed by the National Artificial Intelligence Initiative Act of 2020 (P.L. 116-283). Its stated goal: offer a resource to organizations designing, developing, deploying, or using AI systems to help manage the many risks of AI and promote trustworthy and responsible development and use of AI systems.

That's it. Voluntary. Rights-preserving. Non-sector specific. Use-case agnostic. Provides flexibility to organizations of all sizes and in all sectors.

The framework breaks down into four functions — Govern, Map, Measure, and Manage. Govern is the foundational layer that permeates the others. The Core contains 72 subcategories. The framework lists seven qualities of trustworthy AI: valid and reliable, safe, secure and resilient, accountable and transparent, explainable and interpretable, privacy-enhanced, fair with harmful bias managed.

Note what's absent. The word "misalignment" appears nowhere. Neither does "reward hacking," "deceptive alignment," or "specification gaming." These aren't oversights. They reflect a design choice made in 2022 when the framework was drafted, when the primary AI risk conversation in enterprise settings was about hiring algorithms that encoded gender bias or credit scoring models that produced disparate outcomes.

The Ethics Problem Nobody Wants to Discuss

Neil Raden's assessment on diginomica makes an observation that still stings: run a word count on the AI RMF. "Ethics" appears zero times. "Ethical" appears four times. "Philosophy" gets nothing. Meanwhile, "Govern" shows up 66 times and "Risk" hits 268.

Raden's reading is that the framework treats trustworthiness and responsible AI as a means to an end — the real focus is risk management. Not ethics. Not alignment. Risk.

This distinction matters more than it sounds. Risk management asks: "What's the probability this thing causes harm, and what's the cost?" Alignment asks: "Does this system's objective function actually correspond to what we want?" These are different questions with different tool sets.

If a model is misaligned, genuinely misaligned, not just biased, risk management tells you to monitor it, log its outputs, maybe put a human reviewer in the loop. Alignment would tell you the deployment shouldn't exist at all. The AI RMF has no vocabulary for the second answer.

That's not a criticism of the framework's authors. You can't build what you weren't asked to build. The National AI Initiative Act of 2020 didn't mandate a misalignment framework. It mandated a risk management framework. NIST delivered exactly that.

The Resource Center and the Operationalization Gap

NIST anticipated the "okay, but how?" question. The AI Resource Center (AIRC) is the operationalization layer, the place where the abstract framework meets the practical need for testing, evaluation, verification, and validation (TEVV).

AIRC offers a Playbook with suggested actions and documentation practices. It hosts a Roadmap. There are Example Use Cases, Crosswalk Documents mapping the AI RMF to other frameworks, a Glossary, and an AI Metrology Center for submitting metrics. There's even a Generative AI Profile that extends the AI RMF concepts to specific technologies and sectors.

The technical reports section is telling. NIST published ARIA 0.1 (Assessing Risks and Impacts of AI), which demonstrates a pilot evaluation combining data from expert annotators and human testers with a transparent measurement tool. They're building evaluation infrastructure. There's the Global Engagement on AI Standards (AI 100-5e2025), establishing an engagement plan to promote AI standards based on NIST risk management principles.

All useful. All aimed at a world where the problem is measuring whether a system is performing within acceptable parameters.

What you won't find is evaluation infrastructure for detecting whether a model is strategically compliant during testing but not in deployment. That's a different problem, arguably not solvable with the tools NIST has built, but it's the problem the misalignment reports are pointing at, and there's a silence in the AIRC resources that speaks volumes.

What Enterprises Are Actually Getting

Here's where I'll commit to an opinion: most companies deploying the AI RMF right now are getting a false sense of security. Not because the framework is bad, but because they're pointing it at a threat it doesn't cover.

A mid-size company integrating a foundation model into customer service, running through the Map and Measure functions, checking for bias in outputs and monitoring for drift, that's genuinely valuable work. It addresses real harms. It's not wasted effort.

But if the concern is that a sufficiently capable system might exhibit alignment faking, or that the model's behavior in a specific high-stakes context isn't reliably predicted by its benchmark performance, the AI RMF leaves you with a governance process and a prayer. The framework's answer to "what if the AI is strategically deceptive about its capabilities?" is... there isn't one. Govern the organization. Map the context. Measure outcomes. Manage risks. Fine. Against what? Against a system that knows you're measuring?

Recent lab-to-enterprise disclosures make the concern concrete rather than hypothetical, see our breakdown of OpenAI's model misalignment disclosure framework, our analysis of AI agents passing cover-up instructions to their successors, and our look at model cards revealing foreknowledge of data destruction. These are the kinds of behaviors a risk register was never designed to catch.

This is why I think the misalignment discourse and the enterprise AI governance discourse are currently two parallel conversations that haven't found a good interface yet.

The Revision Cycle and What Comes Next

NIST has flagged that AI RMF 1.0 is being revised. The Playbook will update after the revision. That's the mechanism: the framework evolves, absorbs new concerns, expands its vocabulary.

Whether it can absorb misalignment concerns is an open question. You can add a subcategory about "monitoring for unexpected model behaviors." You can write a playbook entry about red-teaming for alignment failures. You can even build a TEVV methodology that catches some patterns of misaligned behavior at evaluation time.

What you can't do is make a risk management framework think like an alignment researcher, because those are fundamentally different intellectual traditions with different assumptions about what AI systems are.

Enterprises looking for a framework that handles misalignment should treat the AI RMF as necessary but insufficient. Pair it with internal governance structures that address capability thresholds, some of the work we've covered points toward governing agentic AI around operator intent rather than static rules, and with risk transfer mechanisms that price in tail outcomes. The regulatory patchwork in the US is moving, and nobody has solved this cleanly.

The misalignment problem deserves a framework built for it. Until one exists, the AI RMF is the least-bad option for structured thinking about AI risk. Just don't confuse a risk register with an alignment proof. They solve different problems, and pretending otherwise is its own kind of risk.

the misalignment conversation has outgrown its tools

More blogs