ProBackend
ai coding tools endpoint exploitation
2 hours ago5 min read

Ai Driven Endpoint Security Trends: Inside the Heapjack and Overpatch Codex Sandbox Escapes

Security researcher Oren Yomtov discovered two sandbox escape vulnerabilities in OpenAI's Codex (Heapjack and Overpatch) that allowed command execution on developer hosts. OpenAI patched both in August 2026.

As autonomous developer tools sink deeper into corporate infrastructure, safeguarding developer machines has become priority number one. Examining ai driven endpoint security trends reveals a harsh reality: when an AI coding assistant gets a shell on a developer's machine, the sandbox boundary between "assistant" and "attacker" effectively disappears. That boundary just cracked twice in the same product.

Security researcher Oren Yomtov, head of security research at NaaS platform Pentera, disclosed two sandbox escape vulnerabilities in OpenAI's Codex command-line coding assistant that allowed arbitrary command execution directly on the host system. The flaws, which Yomtov dubbed Heapjack and Overpatch, were reported in late June and early July 2026 and patched by OpenAI on August 4. Together they earned a $4,500 bounty. Yomtov's disclosure is titled From Sandboxed Helper to Full Host Compromise with RCE via Escaping OpenAI Codex CLI.

Heapjack: An Out-of-Bounds Write in the Seatbelt Execution Policy

Codex CLI enforces its sandbox at the OS level. On macOS it wraps command execution in sandbox-exec, a kernel-enforced Seatbelt policy, applying one of three profiles:

  • read-only: the most locked-down mode, permitting reads only.
  • workspace-write: read and write access to the working directory, with network access disabled by default.
  • danger-full-access: no sandbox, intended only for trusted environments.

When a rule does not match, Seatbelt falls back to a default action (sandbox-allow or sandbox-deny) that sits in a policy structure called the execution policy. Heapjack, assigned CVE-2026-33238 with a CVSS score of 7.5 (High), is an out-of-bounds write vulnerability in sandbox-exec that operates on exactly that structure.

Codex's Seatbelt profiles contain parameterized regular expressions. Yomtov found that a crafted regex string can write a 32-bit value at a negative offset from the target string, letting an attacker overwrite the execution policy's default action from sandbox-deny to sandbox-allow. The result is that the profile becomes a no-op: every operation the sandbox was supposed to deny is now permitted, including file writes and command execution outside the allowed scope.

Heapjack is triggered when the sandbox evaluates an attacker-controlled path, so the escape works even in read-only mode. Yomtov said the bug was not exclusive to Codex: other AI tools using the same Seatbelt approach, including OpenAI's GPT-5 Codex CLI fork, Amp, and Cursor, were exposed. Google's Gemini CLI was also affected, but with lower impact because it denies the syscalls needed to weaponize the bug; Claude Code, which also uses sandbox-exec, was not affected.

The vulnerability traces back to the original Apple Seatbelt sandbox-exec binary, meaning it ultimately stems from a legacy Apple binary. Because Apple has not updated that component, the root cause cannot be fixed at the OS level, and vendors must mitigate rather than eliminate it.

Overpatch: Breaking the Sandbox with the Patch System Itself

The second issue is a defense mechanism turned against itself. After Heapjack was responsibly disclosed, OpenAI shipped an in-memory binary patch for sandbox-exec that installed a syscall trampoline to block exploitation without waiting for an Apple update. Yomtov's follow-up work turned that patch into a new sandbox escape, which he called Overpatch (no CVE has been assigned).

The binary patch was loaded via DYLD_INSERT_LIBRARIES, with the injected dylib exporting DYLD_INTERPOSE hooks and an __interpose section. Overpatch abuses this by injecting a fake dylib whose interpose entry points to a single ret instruction, causing the patching routine to exit early. Codex believes it is sandboxed; in reality no patch was applied, the sandbox was silently skipped, and the original Heapjack path becomes available again. Yomtov described the trick as "basically patching the patch."

The same technique can also abort patching mid-way, leaving an inconsistent patched binary and creating a new class of vulnerability.

Escape Chain: From Locked-Down Mode to Host Control

Yomtov chained the two issues into a full escape. An attacker first establishes a foothold on the victim machine, then plants malicious content in a project workspace. In his proof-of-concept, simply asking Codex to summarize a text file was enough to detonate the chain, and Codex executed open -a Calculator.app on the host.

The escalation across sandbox modes is what makes this alarming for endpoint teams:

  • workspace-write (Codex's default): arbitrary file write, full read and write access to the working directory and project, and arbitrary command execution on the host.
  • read-only (most restricted): with Overpatch, full read and write across user-writable areas, arbitrary command execution, and the ability to install persistence such as LaunchAgents.

These escapes are not an isolated bug story — they are a preview of where endpoint attacks are heading, a theme echoed in our breakdown of the BragJack browser AI hijacking attacks and in our analysis of process parameter poisoning as an EDR blind spot. AI coding assistants run with developer privileges, touch source code and secrets, and execute commands continuously. When the sandbox that contains them can be silently disabled from the inside, the assistant becomes the attacker's implant.

Three concrete defense actions follow from this research:

  1. Treat the agent as an untrusted execution context. Require confirmation for any command an assistant requests outside the project working directory, and prefer deny-by-default approval policies over full-access modes.
  2. Do not trust vendor-side binary patching as a durable control. In-memory patches like OpenAI's sandbox-exec shim are bypassable (Overpatch proves it) and Apple's legacy Seatbelt binary will not receive a root-cause fix. Layer OS-level controls — EDR telemetry on sandbox-exec invocations, DYLD_INSERT_LIBRARIES abuse detection, and LaunchAgents/persistence monitoring — on top of vendor mitigations.
  3. Instrument dynamic library injection. Overpatch is fundamentally a dylib-injection abuse of DYLD_INTERPOSE. Endpoint detection rules that flag unexpected DYLD_INSERT_LIBRARIES loads against security tooling catch both this class of "patch the patch" attack and the broader EDR-evasion trend.

The Bigger Picture

AI coding assistants are being embedded into enterprise development workflows faster than their containment guarantees have been audited. Heapjack shows a decade-old OS primitive can silently void a sandbox; Overpatch shows that even the vendor's emergency fix can be reversed by the very attacker it is meant to stop. Teams deploying these tools should assume their sandbox is advisory, not binding — and build ai powered endpoint security controls accordingly.

This article is based on the vulnerability disclosure and chain analysis by Oren Yomtov (Pentera), covered by BleepingComputer. As of publication, CVE-2026-33238 (Heapjack, CVSS 7.5) has been published and no CVE has been assigned for Overpatch.

heapjack: an out-of-bounds write in the seatbelt execution

ai driven endpoint security trends in the era of agentic ai systems

More blogs