ProBackend
cloud security incidents
19 hours ago5 min read

ZML's Multi-Chip Inference Server and What It Means for Your Cloud Security Incident Response Playbook

Paris-based ZML, backed by Yann LeCun and $20M in funding, has released ZML/LLMD — a free inference server running LLMs across Nvidia, AMD, Google TPU, Apple Metal and Intel Arc chips. Here's why the vendor lock-in fight matters for security teams.

The Inference Lock-In Problem Nobody Talks About

Here's something most security teams don't think about until it's too late: your AI inference stack is probably locked into a single vendor's silicon. And that creates attack surfaces you haven't audited yet.

ZML, the Paris-based startup endorsed by Turing Award winner Yann LeCun, just released ZML/LLMD — a free inference server that runs open-source large language models across Nvidia GPUs, AMD accelerators, Google TPUs, Apple Metal and Intel Arc. All from one software stack.

Founder Steeve Morin told TechCrunch that inference optimization has been outpacing model training in importance, but the landscape remains patchy. Software and architecture barriers create vendor lock-in that most enterprises accept as inevitable.

I don't think it's inevitable. And I think security teams should care about this more than they do.

When you're locked into one chip vendor, you're locked into their supply chain. Their security posture. Their incident response timeline. If Nvidia has a zero-day in their inference runtime, your entire AI stack goes dark. That's not hypothetical — it's just a matter of when.

ZML is betting that giving enterprises the option to mix chips — including cheaper or lower-energy options — will break that lock-in. Morin says the goal is to "give people back the power to create their own system and achieve real efficiency gains that allow AI to be disseminated."

Translation: stop being held hostage by a single vendor's roadmap.

Why This Matters for Security & Compliance Teams

Let me be direct: heterogeneous compute isn't just an efficiency play. It's a security strategy.

When you can distribute inference workloads across multiple chip architectures, you're not just optimizing costs. You're reducing single points of failure. You're creating options for incident response when one vendor's infrastructure goes down — or gets compromised.

Think about your cloud security incident response playbook. Right now, if your primary inference provider has an outage or a breach, you're scrambling. You don't have a fallback. Your models are stuck on one vendor's hardware, and that vendor controls your recovery timeline.

ZML/LLMD changes that equation. Not dramatically overnight, but it opens the door. Enterprises and cloud providers can now design systems that mix chips — some cheaper, some lower-energy, some from vendors with different security postures.

That doesn't mean you should rip out your Nvidia infrastructure tomorrow. But it does mean you have options. And options are what good incident response is built on.

The Team Behind the Code

Morin brings serious credentials to this. He was VP of engineering at Zenly, the Snapchat-acquired startup that went for nine figures in 2017. That track record helped him raise $20 million from a solid investor group: Harry Stebbings' 20VC, >commit, AALVC, Drysdale Ventures, Xavier Niel's Kima Ventures, Kindred Capital, LocalGlobe and Puzzle Ventures.

The team is lean — 20 people in Paris. Morin credits that small size for their speed. I believe him. Bigger teams move slower. Smaller teams ship faster.

This isn't their first public project either. ZML released an inference-focused ML framework in 2024, updated it in March 2026. They've been working on this problem for a while.

Co-Designing Silicon, Not Just Software

Here's where it gets interesting. Morin says ZML has reached the point of co-designing silicon with chip partners. They're working with European AI chip startups like Axelera, Fractile, Kalray, OLIX, Q.ANT, SiPearl, SpiNNcloud and VSORA.

He's not bearish on Nvidia — ZML has a good relationship with them, and Morin acknowledges their existing supply chain advantages. But the ambition is clear: work on "things that haven't been done before anywhere in the world."

That's a bold claim. And it matters for security teams because silicon co-design means deeper integration. Deeper integration means you can build security controls closer to the hardware. You're not just slapping on software patches — you're designing trust from the ground up.

The Competitive Landscape

The inference space is heating up. TechCrunch calls it the "inference gold rush," and for good reason.

Baseten is valued at $13 billion. Inferact comes from the creators of vLLM. RadixArk sits behind SGLang. Both vLLM and SGLang partially compete with LLMD, but Morin says ZML covers a broader spectrum.

I'd watch this space closely. When you have multiple inference platforms competing, vendors have to innovate faster. They have to prioritize security. They can't rest on their existing customer base.

That's good for everyone, especially security teams who've been asking vendors to take inference security more seriously.

What This Means for Your Playbook

So where does this leave you? If you're a security & compliance analyst reviewing your organization's AI infrastructure, here's what I'd suggest:

First, audit your current inference stack. Where are you locked in? What happens if that vendor goes down for 48 hours? A week?

Second, start evaluating heterogeneous options. You don't have to migrate tomorrow, but you should understand what's possible. ZML/LLMD is one option. Others will emerge.

Third, update your incident response playbook to account for inference-specific failures. Most playbooks focus on compute and storage outages. They don't cover model serving failures, inference runtime vulnerabilities or chip-level supply chain disruptions.

Fourth, talk to your procurement team. Vendor lock-in isn't just a technical problem — it's a contractual one. Make sure your agreements allow for flexibility.

This isn't about switching vendors overnight. It's about recognizing that you have options, and preparing for the day you need to use them.

The Bigger Picture

AI is becoming embedded in our work and everyday lives. That means inference optimization matters more than ever. But it also means the security implications are deeper.

When you're running LLMs on proprietary hardware, you're trusting that vendor with your most sensitive workloads. Their security posture becomes your security posture. Their incident response timeline becomes yours.

ZML/LLMD doesn't solve all of that. But it starts to break the monopoly. And in security, breaking monopolies is how you build resilience.

I'm watching this space. I think heterogeneous inference will become table stakes within the next two years. The question isn't whether you'll adopt it — it's when.

And when that day comes, your cloud security incident response playbook should already be ready.

The Inference Lock-In Problem Nobody Talks About

The Inference Lock-In Problem Nobody Talks About

Sources

Sources

Sources

More blogs