ProBackend
supply chain attacks
1 hour ago5 min read

How Dependency Confusion Turns Build Tools Against Your Infrastructure

An analysis of dependency confusion: how automated package resolution pulls malicious public packages into corporate builds, what Alex Birsan demonstrated across major tech firms, and practical registry safeguards.

When building software, we place a blind kind of trust in package managers. We type npm install or pip install and assume the tooling will grab exactly what we asked for, from wherever we intended. But what happens when your build system gets confused about where a package lives?

Last year, security researcher Alex Birsan showed the industry just how fragile that trust really is. Working from a simple manifest file shared by fellow researcher Justin Gardner, Birsan uncovered a subtle architectural flaw in how popular package managers resolve dependencies. By exploiting what is now known as dependency confusion, he quietly compromised over 35 major tech organizations—including Apple, Microsoft, PayPal, Shopify, Netflix, Tesla, Yelp, and Uber—earning upwards of $130,000 in bug bounties along the way.

The Mechanics of Name Collisions

Most modern development ecosystems rely heavily on public code registries like npm, PyPI, and RubyGems to fetch third-party libraries. At the same time, large engineering teams maintain private registries to host proprietary internal modules.

The problem arises when an internal package shares its name with a public one. If your package manager is configured to check public repositories either before or alongside private ones, it creates an ambiguity. Birsan realized that if an internal dependency name was not claimed on the public registry, anyone could step in and register a public package with that exact same name.

When a developer or an automated CI/CD pipeline subsequently ran an installation command, the package manager would evaluate both sources. In many ecosystems, the public registry takes precedence, or higher version numbers on public feeds automatically win out. Without a single warning, the build system would pull down the attacker-controlled public package instead of the private internal library.

Finding the Blind Spots in Enterprise Code

How did Birsan know what internal names to target? Large companies don't usually publish their private library catalogs out in the open. But leaks happen everywhere if you know where to look.

According to FOSSA’s analysis of the attack vector, internal package names frequently bleed into public view through several common channels:

  • Accidental public code commits on GitHub.
  • Leftover dependency lists (package.json files) embedded directly into public-facing JavaScript build artifacts.
  • Internal error traces, URL paths, and require statements exposed in web applications.

By scouring GitHub repositories and CDN assets, Birsan harvested the names of proprietary packages used by major corporations. Armed with those names, he published blank placeholder packages on npm, PyPI, and RubyGems under his real researcher account, complete with explicit disclaimers stating the packages were for security research purposes only.

Automated Execution and DNS Exfiltration

What made this attack truly potent was its complete lack of friction. Traditional typosquatting relies on developers making a spelling mistake and installing a malicious package manually. Dependency confusion requires zero developer error or social engineering. The build system does the dirty work automatically.

To prove that his public packages were actually entering corporate networks during builds, Birsan embedded simple preinstall scripts into his test packages. Because standard corporate networks are heavily firewalled against inbound connections, he used DNS-based exfiltration.

As soon as a target company's build server pulled in the rogue dependency, the preinstall script fired off a DNS lookup to a domain controlled by the researcher. The incoming query revealed the exact IP address making the request, along with the local username and home directory path of the build environment. Once those callbacks rolled in, Birsan had definitive proof that counterfeit code had successfully executed inside enterprise infrastructure.

Industry Response and Architectural Fixes

The disclosures triggered an immediate wave of remediation across the tech sector. Companies like Yelp patched their internal workflows within hours, while Microsoft awarded Birsan $40,000—their highest bug bounty at the time—and assigned CVE-2021-24105 to the issue in Azure Artifacts.

However, maintainers and platform vendors largely agreed on one crucial point: this was not a traditional bug in the repositories themselves, but an inherent design flaw in how package managers resolve ambiguous names. You cannot simply patch npm or PyPI to prevent developers from naming things whatever they want; the client tooling and enterprise build configurations must be hardened.

Furthermore, as ActiveState notes in their analysis of software supply chain risks, modern engineering teams face relentless pressure to deliver code faster. When alert fatigue sets in and security teams struggle to keep pace with manual reviews, automated upstream defenses become essential. Relying solely on reactive post-deployment triage leaves organizations vulnerable to novel supply chain vectors that bypass traditional vulnerability scanners.

Practical Defenses for Engineering Teams

Securing your build pipeline against dependency confusion requires deliberate configuration changes rather than passive hope. Here are the steps teams are taking right now:

  1. Namespace Reservation: If you use npm, adopt scoped packages. By registering and publishing your internal packages under an organization scope (e.g., @yourcompany/auth), you ensure that nobody outside your organization can squat on that namespace.
  2. Registry Configuration: Configure your package managers and CI/CD runners to strictly isolate private registries. Explicitly disable fallback behavior where the client automatically queries public registries for missing private dependencies.
  3. Internal Placeholder Squatting: For package ecosystems like PyPI that do not support strict organizational namespacing, security teams routinely reserve or squat their internal package names on the public registry with harmless placeholder packages. While not an architectural silver bullet, it blocks opportunistic attackers from grabbing those names first.

Ultimately, dependency confusion serves as a stark reminder that our automated build tools are only as secure as our naming conventions. By eliminating ambiguity in how dependencies are fetched, engineering teams can close off one of the slickest supply chain attack vectors in modern software development.

the mechanics of name collisions

More blogs