ProBackend
ai prompt injection threats
1 hour ago7 min read

Red Lines in the Sandbox: Navigating AI Cybersecurity Threats in Microsoft Copilot

An editorial analysis of Microsoft's pushback against Copilot prompt injection claims, comparing enterprise AI security approaches across AWS Bedrock and Microsoft Copilot.

Red Lines in the Sandbox: Navigating AI Cybersecurity Threats in Microsoft Copilot

Security researchers and enterprise software vendors are locking horns over a fundamental question in modern software architecture: when does an unexpected AI behavior cross the line from an architectural limitation into a genuine security vulnerability?

The debate intensified after a security engineer disclosed multiple issues involving prompt injection and sandbox boundaries in Microsoft's Copilot AI assistant. Microsoft has pushed back against characterizing the reported behaviors as security vulnerabilities. The disagreement reflects a broader challenge for the industry: AI systems can be manipulated through hostile inputs, but determining whether that manipulation constitutes a reportable vulnerability depends on the system's promises, threat model, and impact.

What the researcher reported

The dispute traces to a set of findings published in late 2025 by Emmanouel Levy, a security engineer at Securonix and an MVP, who told BleepingComputer that he identified several anomalous Copilot behaviors. His report described prompt-injection paths that could return data obtained from other users' conversations, leak the contents of a user's OneDrive account, and expose internal Microsoft documentation. Levy also said he could poison Copilot's answer store so that responses influenced later interactions of other users.

Beyond prompt injection, Levy described a sandbox-extension issue: Copilot's web-browsing and search tools run in Microsoft Edge via WebView2, and he reported being able to make Copilot load arbitrary URLs, then inject JavaScript into those pages and execute arbitrary code in the context of the running Edge process. He also found that Copilot's sandbox permits outbound network connections to arbitrary sites. Levy argues that combined with prompt injection, these traits could let an attacker exfiltrate sensitive material retrieved by Copilot during normal work, or make Copilot fetch hostile code that it then runs inside its own environment.

Microsoft's position: "not a security boundary"

Microsoft declined to issue CVEs for any of the reported issues. In a statement, a spokesperson framed Copilot's sandbox strictly: "The Copilot sandbox is not a security boundary and is not designed to prevent arbitrary code execution. Its purpose is to protect users from harmful or disallowed content." The company added that arbitrary code execution inside the sandbox does not create impact against other users, Microsoft services, or the underlying platform. On the OneDrive and internal-documentation disclosure claims, Microsoft said its systems did not detect anomalous behavior.

Microsoft did not dismiss the findings entirely. The company marked the network-egress finding as "significant risk" in its internal triage, said it is weighing mitigations, and noted that tools executing code is documented behavior.

The analogies each side reaches for

Levy's counterargument is an analogy to ChatGPT's sandbox, which Microsoft's own statement referenced: if ChatGPT can block the same behaviors within the same underlying model, he argues, Copilot can too. Microsoft's rebuttal is that its sandbox was purpose-built for consumer safety filtering, while OpenAI's targets code execution, so the two designs are simply different.

Levy also invoked the "Hilbert's hotel" thought experiment to argue that protections which appear robust can still be structurally incomplete. Commentary around the disclosure drew a different historical parallel: prompt injection resembles SQL injection, once dismissed by developers as a nuisance before blocking untrusted input became programming 101. For a generation of developers, GeeksforGeeks-style tutorial material on prompt injection now occupies that same foundational slot — a sign of how quickly the issue has moved from curiosity to baseline hygiene.

AI cybersecurity threats and the vulnerability boundary

Prompt injection can influence an AI model by embedding malicious instructions in content it is asked to process. Indirect prompt injection is especially relevant to assistants that inspect external material, such as documents or web pages. A model may follow instructions embedded in that material rather than treating it solely as data. Whether a particular behavior is a security flaw depends on what the product is designed to protect and what the attack enables; the existence of a prompt injection alone does not establish a breach of a security boundary.

This distinction can create a gap between researchers and vendors. Researchers may emphasize plausible paths from untrusted content to unintended actions, while vendors may assess the same behavior against documented product boundaries, mitigations, and expected model limitations. The practical question is not simply whether the model can be influenced, but whether an attacker can use that influence to cross a meaningful boundary or cause harm beyond the expected operation of the product.

Microsoft's multi-layered defense in detail

In a July 2025 Security Research Center post, Microsoft laid out how it defends Copilot against indirect prompt injection, and that framing underpins its dismissal of Levy's report. The defense stack includes preventative techniques such as hardened system prompts and "spotlighting," a method that isolates untrusted inputs and marks them clearly so the model treats them as data rather than instructions. Detection relies on Microsoft Prompt Shields, integrated with Defender for Cloud so enterprise customers get organization-wide visibility into injection attempts. Mitigation of impact comes from data governance policies, per-user consent workflows before consequential actions, and other controls. Microsoft's stated assumption throughout is that prompt injection cannot be fully prevented — only layered against.

Microsoft has published parallel guidance for developers building on the Microsoft Gateway and MCP-based agent architectures, arguing that injection risk should be managed through permissioning, approval flows, and tool-level isolation rather than expecting the model itself to be injection-proof.

What enterprise AI security approaches can mitigate

Enterprise controls can reduce risk without making prompt injection impossible. Access controls and least privilege limit what an agent can do; human approval can gate consequential actions; isolation and sandboxing can constrain execution; and monitoring can help detect suspicious behavior. These safeguards should be evaluated as layers, not as proof that a model will reliably distinguish trusted instructions from hostile content.

That framing is exactly the one AWS applies to Bedrock Agents. AWS describes guardrails — topical filters, denied-topic policies, PII redaction, and content filters — together with action-group controls as parts of an agent deployment, all intended to constrain what an agent reads and which tools it may invoke. Such controls can help filter some inputs and constrain tool use, but they do not make indirect prompt injection disappear or guarantee that an agent will interpret retrieved content safely. A fair comparison between Bedrock and other enterprise AI approaches, including Microsoft's Copilot stack, therefore asks what data and tools an agent can reach, which actions require approval, how untrusted content is handled and delimited, and what audit and containment mechanisms are available — not merely which vendor claims stronger protection. Notably, the architectural philosophies converge: Bedrock's guardrails and action groups, and Microsoft's spotlighting and Prompt Shields, both assume the model is persuadable and push safety into the surrounding platform.

That convergence matters most as enterprises adopt agentic AI — systems that plan multi-step tasks and act through tools. As MIT Sloan explainers on agentic AI have noted, the moment an assistant can browse, read files, and call APIs on a user's behalf, the blast radius of a successful injection grows from a bad answer to a real action taken with the user's credentials.

A practical way to assess AI prompt security

Organizations assessing AI cybersecurity threats should define the assets and trust boundaries first, then test realistic attack paths. Useful questions include whether untrusted input can access sensitive data, trigger tools or actions, alter system behavior across sessions, or bypass approval and isolation controls. Findings should be reported with reproducible evidence and impact, while vendors should clearly explain the intended security boundary and available mitigations.

The Copilot dispute is a reminder that AI security evaluation must account for both model behavior and product architecture. Clear threat models, constrained permissions, careful human oversight, and transparent vulnerability handling help distinguish an inherent limitation from an exploitable security issue — and reduce risk either way. Even where vendors and researchers never agree on the label, enterprises that treat prompt injection as an operational threat to be engineered around, rather than a bug to be awaited, will be on the right side of the argument in 2026.

Related reading: AI cybersecurity threats in 2026: ASCII smuggling slips past filters and into your Copilot.

red lines in the sandbox

More blogs