ProBackend
model performance safety
2 hours ago6 min read

GLM-5.2 and the Open-Weight Safety Black Hole

A new SaferAI report finds Z.ai's open-weight GLM-5.2 approaches frontier AI capabilities while lacking key safety mitigations, renewing concerns that powerful open models could outpace governance and regulatory oversight.

The Open-Weight Paradox

A Chinese open-weight model has quietly closed the gap with the world's most powerful AI systems. The catch? It has almost none of the guardrails those systems carry.

GLM-5.2, developed by Beijing-based Z.ai, sits just months behind OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7 on benchmarks measuring cybersecurity and biological research capabilities, according to a new report from the France-based AI safety nonprofit SaferAI. The capability gap is narrowing fast. The safety gap, meanwhile, is widening into something resembling a chasm.

Zero Refusals on Offensive Tasks

The numbers from SaferAI's evaluation are stark. GLM-5.2 refused none of the offensive cybersecurity or biology tasks it was given through Z.ai's public API.

Compare that to Claude Opus 4.7, which "refused so consistently that SaferAI could not complete CyberGym on it at all," the report notes. CyberGym is a benchmark designed to evaluate a model's cybersecurity capabilities—and its propensity to refuse dangerous requests.

OpenAI used the same benchmark in its evaluation leading up to last month's Hugging Face breach. The contrast between GLM-5.2's total lack of refusals and Claude's near-total refusal rate is impossible to ignore. It's the difference between a model that helps you and one that won't.

The Unenforceable Safety Problem

Here's what makes open-weight models so uniquely problematic. Z.ai could apply safety measures to its hosted API—maybe it already does. But once someone downloads the weights and runs them on their own hardware, those protections become unenforceable.

Anyone can modify the safeguards. Fine-tune the model. Change the system prompts. There's no API-level control to fall back on, no classifier to invoke, no refusal training that holds when the model lives on a server you control.

Frontier developers like OpenAI and Anthropic rely on layers of safeguards—classifiers, refusal training, API-level controls—to limit dangerous cyber and biological assistance. Those measures aren't perfect. Jailbreaks routinely bypass protections on deployed models. AI safety nonprofit Far.ai recently found hundreds of universal jailbreaks—reusable keys that succeed across most harmful requests, in frontier models like xAI's Grok 4.5 and Google DeepMind's Gemini 3.1 Pro.

But none of these safeguards work at all on open-weight models. By design, open-weight models run on any infrastructure, with any set of safeguards, or without any at all.

What Z.ai Didn't Publish (and Wouldn't)

SaferAI's report found that Z.ai didn't publish a safety framework, pre-deployment testing commitments, or risk assessment for GLM-5.2. TechCrunch asked Z.ai whether it conducted internal or third-party frontier safety evaluations before release. The company didn't respond.

That silence matters. It leaves the question of whether GLM-5.2 was intentionally designed to lack safety mitigations, or whether the absence is simply a byproduct of Z.ai's release priorities, entirely unanswered.

Henry Papadatos, SaferAI's executive director, put it plainly: "The frontier of capability is not the frontier of risk, and so we do have to take into account the state of the mitigations as well to assess the risk properly."

The Data Filtering Dead End

One technique Papadatos flagged could help: pre-training data filtering. The idea is simple enough, remove offensive cybersecurity information from training data, then train the model on the curated dataset.

Some research suggests this can reduce hazardous biological knowledge without harming overall model performance. For cybersecurity? Not so much.

It's difficult to train a general model that excels at coding but isn't also good at hacking. Coding has become AI's biggest moneymaker. Developers face enormous pressure to keep improving those capabilities, even as they search for ways to limit misuse. The market doesn't reward models that can't code.

So frontier developers have increasingly relied on other mitigations instead. Anthropic's Opus 5, for example, can search for vulnerabilities in uncompiled source code, but not compiled software, per the model's system card. The reasoning is straightforward: this makes it harder to use Opus 5 for offensive purposes.

Others include rigorous pre-deployment safety evaluations, publishing risk assessments, and, when the model is deemed too dangerous with too many safeguards, simply withholding the weights entirely.

GLM-5.2's release, as far as SaferAI can tell, involves none of these.

The Chinese Policy Context

Chinese leaders have increasingly acknowledged the risks of advanced AI. At last month's World AI Conference, Chinese President Xi Jinping emphasized the importance of open-weight models while also stressing the necessity of ensuring AI remains a tool under strict human control.

Graham Webster, who studies Chinese AI policy at the Stanford Cyber Policy Center, told TechCrunch that China has robust regulations governing AI. But those rules have historically focused on politically sensitive content, misinformation, and social stability, not catastrophic AI risks like offensive cyber capabilities or biological misuse.

"U.S. AI thinkers are, in general, more concerned with this existential catastrophic [idea] than the Chinese community," Webster said. Many Chinese policy researchers, he added, believe that if there's truly going to be a novel frontier risk, American companies will likely encounter it first.

"The Chinese system has confidence that they control the use of these technologies inside China," Webster continued. "Being online in China is something you do attributed to your real name, and companies can be held accountable, users can be held accountable."

Webster mused that the same mechanism model providers use for refusing to engage on certain political topics could potentially be tweaked to make sure models refuse to complete offensive cyber attacks or deliver adverse biological engineering outcomes. But because Chinese companies tend to coordinate with regulators behind the scenes, it's tough to know what internal testing they're conducting before release.

The Open-Source Defense Argument

Advocates of open-weight AI argue that releasing weights is critical for cybersecurity. It allows companies to defend themselves against attacks. Hugging Face relied on GLM-5.2 to defend itself against OpenAI's breach, for example. It also allows organizations to better prepare for future threats if they know what's coming.

"The same systems that helped stop an AI-powered cyberattack can now help defend against millions of cyberattacks every day, while helping us identify and fix vulnerabilities before attackers exploit them," Hugging Face CEO Clem Delangue said in a social media post this week.

Papadats isn't buying it. The benefit, he argues, is often overstated. It doesn't mean "we should open-source dangerous capabilities."

"The main point in my mind is that we shouldn't just accept that dangerous capabilities are easily accessible by anyone anywhere," he said.

His reasoning is grounded in a simple asymmetry: attackers adopt new tools faster than defenders do. A ransomware group can change its methods in a week. A hospital cannot.

That gap is where the risk lives. And GLM-5.2, with its frontier-level capabilities and zero safety mitigations, is the widest expression of it yet.

What Comes Next

The debate over open-weight AI is shifting. It's no longer about whether these models can compete with frontier closed systems. They already can. The question is how society manages the risks once these models are released without the safety infrastructure that accompanies proprietary alternatives.

The answer, currently, is: not well.

GLM-5.2 represents a turning point. It's the first open-weight model to narrow the gap with the world's most capable AI systems while entirely bypassing the safety frameworks those systems carry. Whether that's a feature or a bug depends on who you ask. But the risk it creates is real, and it's only going to grow as more models follow this pattern.

The frontier of capability has moved. The frontier of safety hasn't. That's the paradox, and the problem, of open-weight AI in 2026.

the open-weight paradox

More blogs