ProBackend
agentic ai security risks
2 days ago7 min read

The Stack Behind the Kill Chain

An investigation into how AI systems automate six stages of military targeting in modern warfare, examining the Maven Smart System, Operation Epic Fury, and the inherent flaws of machine learning in life-and-death decisions.

The Stack Behind the Kill Chain

You don't need to believe in killer robots to worry about the stack of AI systems now mediating modern combat. The real story isn't about sci-fi dystopias—it's about mundane integration failures, automation bias, and the way machine learning quietly erodes human oversight across six distinct stages of military targeting.

Airwars, a not-for-profit transparency watchdog, published a detailed report last week titled Anatomy of an AI Kill Chain. Co-authored by Sophia Goodfriend, Heidy Khlaaf, Namir Shabibi, Joe Dyke, and Nathan Walker, the investigation offers a visual guide through a fictitious kill chain—the steps real militaries go through to identify targets and eliminate them—with an emphasis on decisions delegated to AI. It's a rare break in military secrecy, and it matters.

Six Stages, Two Humans

A recent book about US military AI revealed something unsettling: in some operations, only two of the six stages of the US military kill chain now involve humans in the loop. A third involves partial human oversight. The remaining four are fully automated.

The six stages are: decision support systems for data gathering, surveillance technology, intelligence and identification, target selection, strikes on targets, and post-strike assessments.

Sophia Goodfriend, research fellow at the University of Cambridge's Pembroke College and co-author of the report, puts it bluntly. "The stack of AI systems integrated into warfare are the subject of a lot of hype and a lot of mystique," she told The Register. "It's rare for those making or deploying these systems to really break down how they actually operate."

Her point isn't that these are teething problems that will be mitigated once technologies are deployed more and refined more. It's that machine learning systems make errors—inherent, predictable errors that can't be engineered away.

Decision Support: The Maven Smart System

The Pentagon's Maven Smart System, produced by Palantir Technologies, collects information on enemy positions from radar signals, satellite and drone imagery, electronic communications, and other sources. It combines everything into a "common operating picture" of the battlefield.

Cameron Stanley, the Pentagon's chief digital and AI officer, told Palantir's AIPCon 9 conference on March 12 that Maven fuses everything "into a single visualization tool" for use in decision-making. Instead of consulting eight or nine separate systems, commanders now identify a target and select a strike package by clicking on one screen.

"From identifying the target to now coming up with a course of action to now actioning that target, all from one system. This is revolutionary," Stanley said.

Project Maven started in 2017 under the Pentagon's Algorithmic Warfare Cross-Functional Team (later the Joint Artificial Intelligence Center). It was designed to use computer vision and machine learning to sort through drone footage of Middle Eastern battlegrounds and identify potential militant hideouts. Over time, Palantir steadily improved the technology, enabling it to collect and collate data from multiple sources and to recommend possible combat moves—including identifying the friendly unit best positioned to strike any given target, and generating courses of action with attack vectors and munitions to be employed.

In late 2024, Anthropic's Claude AI operating system was merged with Maven to provide enhanced targeting options. Commanders can now generate target lists by criteria such as radar station, missile battery, communications node, or senior commander, rank them by strategic importance, and once a target has been attacked, the system reviews damage assessment reports and automatically produces new target lists—all in a matter of minutes.

Anthropic has since been barred from providing services to the US military due to its refusal to work on autonomous weapons systems. Other AI companies, including OpenAI, have been recruited to assume Anthropic's role.

For military officials, Maven's speed in selecting and reselecting targets represents a distinct combat advantage. "Our war fighters are leveraging a variety of advanced AI tools," said Adm. Brad Cooper, commander of US Central Command, in a March 11 video briefing. "These systems help us sift through vast amounts of data in seconds so our leaders can cut through the noise and make smarter decisions faster than the enemy can react."

The Iran Strikes: A Case Study

Operation Epic Fury—the US air and missile campaign against Iran that began February 28—offers a concrete example of what happens when AI-mediated targeting goes wrong. According to an April 8 White House accounting, the US military struck more than 13,000 targets in Iran during the first 38 days of the war: more than 2,000 command-and-control targets, 1,500 air defense targets, and 1,450 industrial base targets.

The White House claimed all were legitimate military targets. The New York Times and other news media reported the visual destruction of civilian facilities.

The Shajareh Tayyebeh girls' elementary school in Minab, Iran, was struck by a US cruise missile on February 28, killing over 170 people, most of them children. According to a preliminary assessment by Central Command, US intelligence maps failed to indicate that the school facility—once part of a military base—had long ago been converted to civilian use. It had been added to an AI-generated target list without adequate human supervision.

"The New York Times reported March 11 that 'officials said the error was unlikely to have been the result of new technology' but rather 'human error in wartime,'" The National, an Abu Dhabi-based newspaper, reported. But the distinction between human error and automation bias is thinner than officials admit.

Automation Bias: The Real Enemy

Nilza Amaral, head of research at Chatham House's Global Governance and Security Centre, told The National there's "a concern that targeting [approval] could end up just being a mere formality because of the automation bias, where people are just relying on what the machine is telling them."

Automation bias—the tendency for humans to defer to automated outputs even when they should verify them—is the pernicious problem Goodfriend identifies. Under wartime time constraints, it's tempting and much easier to defer to an AI-generated output than to double-check translations, assessments, or target lists.

Even if you have somebody reviewing a risk profiling system by looking at a social media post, that social media post will be translated by a machine learning algorithm. You won't be able to verify if that translation is correct because you're dealing with time constraints. It will be really tempting and much easier to defer to an AI-generated output.

With humans granted ever-diminishing time in which to review AI-derived targeting decisions, the risk of error naturally increases, many analysts say. Whether AI played a role in the Shajareh Tayyebeh strike is unclear—U.S. intelligence maps simply failed to reflect that the facility had been converted to civilian use. But the growing danger is that "humans may rely too much on the system" and fail to double-check its recommendations, Amaral noted.

The Inherent Flaws of Machine Learning

The Airwars report touches on sources of potential errors at every stage: the reliability and accuracy of decision support systems, automated translation errors during surveillance of a target's text messages, target risk scoring systems derived from algorithmic assessment of social media data, computer vision systems that may misidentify objects and people, the shortcomings of recurrent neural networks used during drone strikes when there's GPS and/or electronic jamming, and the potential problems with AI-assisted battle damage assessment systems.

Goodfriend emphasizes these aren't problems that will disappear with more deployment. "The point is not that these are teething problems that will be mitigated or ironed out once these technologies are deployed more and refined more, but really to emphasize the inherent flaws of machine learning and how they can't perform the tasks that militaries expect them to perform when they're deployed on the battlefield."

The specific kill chain described is fictitious, but Goodfriend said that was necessary to avoid making overreaching claims about how specific militaries operate. "Obviously it's hard to reconstruct proprietary systems used by militaries in real time today because they are often censored and secret and reporters cannot pry open exactly how a military is operating."

What Comes Next

The report's goal is to look beyond specific technical systems at the way warfare is mediated by machine learning algorithms and to highlight the specific limitations of those systems. Goodfriend acknowledges it's difficult for civilians to make decisions about technologies they use that might amplify physical safety risks from military AI systems. That's why she says it's important to draw attention to the way these systems work.

"By illustrating how these technologies work, the massive amounts of data, of people's private information, of passive surveillance that is the foundation and lifeblood of automated warfare, I think that that's one way of moving the needle in terms of advocacy," she said.

The stack of AI systems integrated into warfare is real. It's not coming. It's here. And until we demand transparency about how it operates, civilians will keep paying the price.

More blogs