
While policymakers are in discussion on how to regulate the ever-advancing AI technologies such as OpenAI’s GPT-5.6 Sol and Anthropic’s Mythos, a Chinese model with open weights has significantly closed the gap with the leading players in the industry.
The GLM-5.2 model, originating from China’s Z.ai and utilizing open weights, is reported to be just a few months behind OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 in terms of cyber and bio capabilities, based on findings from AI safety nonprofit SaferAI. However, the gap between cutting-edge capabilities and safety measures is expanding.
SaferAI’s analysis, carried out via Z.ai’s accessible API, indicates that GLM-5.2 did not decline any of the requested offensive cyber or dual-use biology assignments. In contrast, Claude Opus 4.7 allegedly “refused so consistently that SaferAI could not finish CyberGym with it.” (CyberGym serves as a standard for assessing cybersecurity capabilities. OpenAI incorporated it in their evaluation following last month’s Hugging Face breach.)
This serves as a stark reminder of concerns raised by some critics over the years: that open-weight AI models could enable highly capable AI to fall into the hands of potential aggressors, with no means to oversee their usage of the technology once they retrieve the weights. As open-weight models rapidly close in on the capabilities of the top AI systems globally, the focus of the debate is shifting from whether they can compete to how society can mitigate risks once they are deployed.
“The cutting-edge of capability does not equate to the cutting-edge of risk, thus it is essential to consider the efficacy of the mitigations to accurately gauge risk,” Henry Papadatos, executive director of SaferAI, shared with TechCrunch.
Although Z.ai might implement safety protocols in its hosted API, these protections turn unenforceable once individuals operate the weights on their own systems, where they can eliminate or modify any safeguards, adjust the models, or alter system prompts.
Developers at the forefront, such as OpenAI and Anthropic, generally rely on safety measures like classifiers, refusal training, and API-level restrictions to limit hazardous cyber and biological support.
However, these measures aren’t infallible: jailbreaks frequently circumvent protections on operational models. Far.ai, an AI safety nonprofit, identified hundreds of universal jailbreaks — marked as reusable keys that succeed on the majority of harmful requests — in advanced models such as xAI’s Grok 4.5 and Google DeepMind’s Gemini 3.1 Pro. The report states that jailbreaks succeed when attackers amalgamate various manipulation tactics — including roleplaying, authority impersonation, false conversation histories, and follow-up prompts — to exploit vulnerabilities in a model’s defenses.
Yet, the safeguards established for closed models are ineffective against open-weight models, which are crafted to function on any infrastructure with any array of safeguards — or none at all.
“The aim should be to ensure that beneficial capabilities — the secure ones — are available to everyone, while attempting to eliminate the harmful ones, even in an open-source manner,” Papadatos stated.
Papadatos mentioned that a possible beneficial technique is “pre-training data filtering,” which entails an AI company removing harmful cybersecurity data from its training sets and then training the model on the sanitized dataset.
Some studies indicate this method can minimize hazardous biological knowledge without compromising overall model efficiency. However, in cybersecurity, data filtering proves to be far less feasible.
It is challenging to develop a general model that excels at programming yet doesn’t also demonstrate talent as a hacker. With coding becoming the most lucrative domain for AI, developers encounter pressure to continue enhancing these abilities even as they seek ways to curtail misuse.
Consequently, leading developers have increasingly turned to alternative mitigations instead. One strategy involves selectively limiting the types of cybersecurity support models are authorized to offer. For instance, Anthropic’s Opus 5 can identify vulnerabilities in uncompiled source code but not in compiled software, as per the model’s system documentation. The rationale is that this limitation makes it more difficult to exploit Opus 5 for offensive objectives.
Other strategies entail thorough pre-deployment safety assessments, publishing risk evaluations, and withholding model weights whenever a system is deemed excessively dangerous.
In the instance of GLM-5.2, SaferAI claims that Z.ai did not disclose a safety framework, pre-deployment testing pledges, or risk evaluation for the model. TechCrunch inquired about whether Z.ai conducted internal or external assessments for frontier safety prior to release but did not receive a reply.
Chinese authorities have increasingly recognized the potential dangers of advanced AI. At last month’s World AI Conference, Chinese President Xi Jinping underscored the significance of open-weight models while highlighting the necessity of ensuring AI remains a tool under strict human oversight.
Graham Webster, who examines Chinese AI policy at the Stanford Cyber Policy Center, conveyed to TechCrunch that China has stringent regulations addressing AI, yet these guidelines have traditionally concentrated on politically sensitive subject matter, misinformation, and societal stability rather than catastrophic AI threats such as offensive cyber abilities and biological misuse.
“Generally, U.S. AI experts are more focused on this existential catastrophic notion than the Chinese community,” Webster remarked, adding that many researchers in Chinese policy expect American companies will likely confront any genuinely novel frontier risk first.
“The Chinese framework is confident they manage the application of these technologies within China,” Webster continued. “Being online in China is linked to your real identity, and firms and users can be held liable.”
Webster speculated that the same mechanisms model providers employ to refuse engagement on certain political subjects may be adapted to ensure models decline to execute offensive cyber operations or provide adverse biological engineering results. He noted that since Chinese firms typically coordinate with regulators behind closed doors, it can be challenging to ascertain what internal evaluations they perform prior to public release.
Proponents of open-weight AI contend that disclosing the weights is vital for cybersecurity because it enables companies to protect themselves against assaults — Hugging Face utilized GLM-5.2 to defend against OpenAI’s breach — and ensures they are better equipped to confront future threats by being aware of what is anticipated.
“The identical systems that aided in thwarting an AI-driven cyberattack can now help to defend against countless cyber threats daily while assisting us in identifying and rectifying vulnerabilities before they can be exploited by attackers,” Clem Delangue, CEO of Hugging Face, stated this week in a social media update.
Papadatos remarked that this advantage is often exaggerated and doesn’t imply “we should make dangerous capabilities open source.”
“The critical point for me is that we must not accept that dangerous capabilities are readily accessible to anyone, anywhere,” he emphasized, asserting that the industry should aim to restrict access to only “good capabilities.” Attackers typically adapt to new tools at a faster rate than defenders. For instance, a ransomware group can modify its tactics within a week, while a hospital cannot.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

