{"id":3491523,"date":"2026-08-04T20:05:26","date_gmt":"2026-08-04T20:05:26","guid":{"rendered":"https:\/\/techingeek.com\/index.php\/2026\/08\/04\/open-weight-ai-models-are-closing-in-on-the-cutting-edge-the-safety-gap-persists\/"},"modified":"2026-08-04T20:05:26","modified_gmt":"2026-08-04T20:05:26","slug":"open-weight-ai-models-are-closing-in-on-the-cutting-edge-the-safety-gap-persists","status":"publish","type":"post","link":"https:\/\/techingeek.com\/index.php\/2026\/08\/04\/open-weight-ai-models-are-closing-in-on-the-cutting-edge-the-safety-gap-persists\/","title":{"rendered":"Open-weight AI models are closing in on the cutting edge. The safety gap persists."},"content":{"rendered":"<div><img decoding=\"async\" src=\"https:\/\/techingeek.com\/wp-content\/uploads\/2026\/08\/open-weight-ai-models-are-closing-in-on-the-cutting-edge-the-safety-gap-persists.png\" class=\"ff-og-image-inserted\"><\/div>\n<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">While policymakers are in discussion on how to regulate the ever-advancing AI technologies such as OpenAI\u2019s GPT-5.6 Sol and Anthropic\u2019s Mythos, a Chinese model with open weights has significantly closed the gap with the leading players in the industry.\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The GLM-5.2 model, originating from China&#8217;s Z.ai and utilizing open weights, is reported to be just a few months behind OpenAI\u2019s GPT-5.5 and Anthropic\u2019s Claude Opus 4.7 in terms of cyber and bio capabilities, based on findings from AI safety nonprofit SaferAI. However, the gap between cutting-edge capabilities and safety measures is expanding.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">SaferAI&#8217;s analysis, carried out via Z.ai\u2019s accessible API, indicates that GLM-5.2 did not decline any of the requested offensive cyber or dual-use biology assignments. In contrast, Claude Opus 4.7 allegedly \u201crefused so consistently that SaferAI could not finish CyberGym with it.\u201d (CyberGym serves as a standard for assessing cybersecurity capabilities. OpenAI incorporated it in their evaluation following last month\u2019s Hugging Face breach.)<\/p>\n<p class=\"wp-block-paragraph\">This serves as a stark reminder of concerns raised by some critics over the years: that open-weight AI models could enable highly capable AI to fall into the hands of potential aggressors, with no means to oversee their usage of the technology once they retrieve the weights. As open-weight models rapidly close in on the capabilities of the top AI systems globally, the focus of the debate is shifting from whether they can compete to how society can mitigate risks once they are deployed.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe cutting-edge of capability does not equate to the cutting-edge of risk, thus it is essential to consider the efficacy of the mitigations to accurately gauge risk,\u201d Henry Papadatos, executive director of SaferAI, shared with TechCrunch.<\/p>\n<p class=\"wp-block-paragraph\">Although Z.ai might implement safety protocols in its hosted API, these protections turn unenforceable once individuals operate the weights on their own systems, where they can eliminate or modify any safeguards, adjust the models, or alter system prompts.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Developers at the forefront, such as OpenAI and Anthropic, generally rely on safety measures like classifiers, refusal training, and API-level restrictions to limit hazardous cyber and biological support.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">However, these measures aren&#8217;t infallible: jailbreaks frequently circumvent protections on operational models. Far.ai, an AI safety nonprofit, identified hundreds of universal jailbreaks \u2014 marked as reusable keys that succeed on the majority of harmful requests \u2014 in advanced models such as xAI\u2019s Grok 4.5 and Google DeepMind\u2019s Gemini 3.1 Pro. The report states that jailbreaks succeed when attackers amalgamate various manipulation tactics \u2014 including roleplaying, authority impersonation, false conversation histories, and follow-up prompts \u2014 to exploit vulnerabilities in a model&#8217;s defenses.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Yet, the safeguards established for closed models are ineffective against open-weight models, which are crafted to function on any infrastructure with any array of safeguards \u2014 or none at all.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe aim should be to ensure that beneficial capabilities \u2014 the secure ones \u2014 are available to everyone, while attempting to eliminate the harmful ones, even in an open-source manner,\u201d Papadatos stated.<\/p>\n<p class=\"wp-block-paragraph\">Papadatos mentioned that a possible beneficial technique is &#8220;pre-training data filtering,&#8221; which entails an AI company removing harmful cybersecurity data from its training sets and then training the model on the sanitized dataset.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Some studies indicate this method can minimize hazardous biological knowledge without compromising overall model efficiency. However, in cybersecurity, data filtering proves to be far less feasible.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">It is challenging to develop a general model that excels at programming yet doesn\u2019t also demonstrate talent as a hacker. With coding becoming the most lucrative domain for AI, developers encounter pressure to continue enhancing these abilities even as they seek ways to curtail misuse.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Consequently, leading developers have increasingly turned to alternative mitigations instead. One strategy involves selectively limiting the types of cybersecurity support models are authorized to offer. For instance, Anthropic\u2019s Opus 5 can identify vulnerabilities in uncompiled source code but not in compiled software, as per the model\u2019s system documentation. The rationale is that this limitation makes it more difficult to exploit Opus 5 for offensive objectives.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Other strategies entail thorough pre-deployment safety assessments, publishing risk evaluations, and withholding model weights whenever a system is deemed excessively dangerous.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In the instance of GLM-5.2, SaferAI claims that Z.ai did not disclose a safety framework, pre-deployment testing pledges, or risk evaluation for the model. TechCrunch inquired about whether Z.ai conducted internal or external assessments for frontier safety prior to release but did not receive a reply.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Chinese authorities have increasingly recognized the potential dangers of advanced AI. At last month\u2019s World AI Conference, Chinese President Xi Jinping underscored the significance of open-weight models while highlighting the necessity of ensuring AI remains a tool under strict human oversight.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Graham Webster, who examines Chinese AI policy at the Stanford Cyber Policy Center, conveyed to TechCrunch that China has stringent regulations addressing AI, yet these guidelines have traditionally concentrated on politically sensitive subject matter, misinformation, and societal stability rather than catastrophic AI threats such as offensive cyber abilities and biological misuse.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cGenerally, U.S. AI experts are more focused on this existential catastrophic notion than the Chinese community,\u201d Webster remarked, adding that many researchers in Chinese policy expect American companies will likely confront any genuinely novel frontier risk first.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe Chinese framework is confident they manage the application of these technologies within China,\u201d Webster continued. \u201cBeing online in China is linked to your real identity, and firms and users can be held liable.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Webster speculated that the same mechanisms model providers employ to refuse engagement on certain political subjects may be adapted to ensure models decline to execute offensive cyber operations or provide adverse biological engineering results. He noted that since Chinese firms typically coordinate with regulators behind closed doors, it can be challenging to ascertain what internal evaluations they perform prior to public release.<\/p>\n<p class=\"wp-block-paragraph\">Proponents of open-weight AI contend that disclosing the weights is vital for cybersecurity because it enables companies to protect themselves against assaults \u2014 Hugging Face utilized GLM-5.2 to defend against OpenAI\u2019s breach \u2014 and ensures they are better equipped to confront future threats by being aware of what is anticipated.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe identical systems that aided in thwarting an AI-driven cyberattack can now help to defend against countless cyber threats daily while assisting us in identifying and rectifying vulnerabilities before they can be exploited by attackers,\u201d Clem Delangue, CEO of Hugging Face, stated this week in a social media update.<\/p>\n<p class=\"wp-block-paragraph\">Papadatos remarked that this advantage is often exaggerated and doesn&#8217;t imply \u201cwe should make dangerous capabilities open source.\u201d <\/p>\n<p class=\"wp-block-paragraph\">\u201cThe critical point for me is that we must not accept that dangerous capabilities are readily accessible to anyone, anywhere,\u201d he emphasized, asserting that the industry should aim to restrict access to only \u201cgood capabilities.\u201d Attackers typically adapt to new tools at a faster rate than defenders. For instance, a ransomware group can modify its tactics within a week, while a hospital cannot.\u201d<\/p>\n<\/div>\n<p><em>When you purchase through links in our articles, we may earn a small commission. This doesn\u2019t affect our editorial independence.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<div><img decoding=\"async\" src=\"https:\/\/techingeek.com\/wp-content\/uploads\/2026\/08\/open-weight-ai-models-are-closing-in-on-the-cutting-edge-the-safety-gap-persists.png\" class=\"ff-og-image-inserted\"><\/div>\n<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">While policymakers are in discussion on how to regulate the ever-advancing AI technologies such as OpenAI\u2019s GPT-5.6 Sol and Anthropic\u2019s Mythos, a Chinese model with open weights has significantly closed the gap with the leading players in the industry.\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The GLM-5.2 model, originating from China&#8217;s Z.ai and utilizing open weights, is reported to be just a few months behind OpenAI\u2019s GPT-5.5 and Anthropic\u2019s Claude Opus 4.7 in terms of cyber and bio capabilities, based on findings from AI safety nonprofit SaferAI. However, the gap between cutting-edge capabilities and safety measures is expanding.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">SaferAI&#8217;s analysis, carried out via Z.ai\u2019s accessible API, indicates that GLM-5.2 did not decline any of the requested offensive cyber or dual-use biology assignments. In contrast, Claude Opus 4.7 allegedly \u201crefused so consistently that SaferAI could not finish CyberGym with it.\u201d (CyberGym serves as a standard for assessing cybersecurity capabilities. OpenAI incorporated it in their evaluation following last month\u2019s Hugging Face breach.)<\/p>\n<p class=\"wp-block-paragraph\">This serves as a stark reminder of concerns raised by some critics over the years: that open-weight AI models could enable highly capable AI to fall into the hands of potential aggressors, with no means to oversee their usage of the technology once they retrieve the weights. As open-weight models rapidly close in on the capabilities of the top AI systems globally, the focus of the debate is shifting from whether they can compete to how society can mitigate risks once they are deployed.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe cutting-edge of capability does not equate to the cutting-edge of risk, thus it is essential to consider the efficacy of the mitigations to accurately gauge risk,\u201d Henry Papadatos, executive director of SaferAI, shared with TechCrunch.<\/p>\n<p class=\"wp-block-paragraph\">Although Z.ai might implement safety protocols in its hosted API, these protections turn unenforceable once individuals operate the weights on their own systems, where they can eliminate or modify any safeguards, adjust the models, or alter system prompts.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Developers at the forefront, such as OpenAI and Anthropic, generally rely on safety measures like classifiers, refusal training, and API-level restrictions to limit hazardous cyber and biological support.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">However, these measures aren&#8217;t infallible: jailbreaks frequently circumvent protections on operational models. Far.ai, an AI safety nonprofit, identified hundreds of universal jailbreaks \u2014 marked as reusable keys that succeed on the majority of harmful requests \u2014 in advanced models such as xAI\u2019s Grok 4.5 and Google DeepMind\u2019s Gemini 3.1 Pro. The report states that jailbreaks succeed when attackers amalgamate various manipulation tactics \u2014 including roleplaying, authority impersonation, false conversation histories, and follow-up prompts \u2014 to exploit vulnerabilities in a model&#8217;s defenses.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Yet, the safeguards established for closed models are ineffective against open-weight models, which are crafted to function on any infrastructure with any array of safeguards \u2014 or none at all.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe aim should be to ensure that beneficial capabilities \u2014 the secure ones \u2014 are available to everyone, while attempting to eliminate the harmful ones, even in an open-source manner,\u201d Papadatos stated.<\/p>\n<p class=\"wp-block-paragraph\">Papadatos mentioned that a possible beneficial technique is &#8220;pre-training data filtering,&#8221; which entails an AI company removing harmful cybersecurity data from its training sets and then training the model on the sanitized dataset.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Some studies indicate this method can minimize hazardous biological knowledge without compromising overall model efficiency. However, in cybersecurity, data filtering proves to be far less feasible.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">It is challenging to develop a general model that excels at programming yet doesn\u2019t also demonstrate talent as a hacker. With coding becoming the most lucrative domain for AI, developers encounter pressure to continue enhancing these abilities even as they seek ways to curtail misuse.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Consequently, leading developers have increasingly turned to alternative mitigations instead. One strategy involves selectively limiting the types of cybersecurity support models are authorized to offer. For instance, Anthropic\u2019s Opus 5 can identify vulnerabilities in uncompiled source code but not in compiled software, as per the model\u2019s system documentation. The rationale is that this limitation makes it more difficult to exploit Opus 5 for offensive objectives.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Other strategies entail thorough pre-deployment safety assessments, publishing risk evaluations, and withholding model weights whenever a system is deemed excessively dangerous.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In the instance of GLM-5.2, SaferAI claims that Z.ai did not disclose a safety framework, pre-deployment testing pledges, or risk evaluation for the model. TechCrunch inquired about whether Z.ai conducted internal or external assessments for frontier safety prior to release but did not receive a reply.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Chinese authorities have increasingly recognized the potential dangers of advanced AI. At last month\u2019s World AI Conference, Chinese President Xi Jinping underscored the significance of open-weight models while highlighting the necessity of ensuring AI remains a tool under strict human oversight.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Graham Webster, who examines Chinese AI policy at the Stanford Cyber Policy Center, conveyed to TechCrunch that China has stringent regulations addressing AI, yet these guidelines have traditionally concentrated on politically sensitive subject matter, misinformation, and societal stability rather than catastrophic AI threats such as offensive cyber abilities and biological misuse.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cGenerally, U.S. AI experts are more focused on this existential catastrophic notion than the Chinese community,\u201d Webster remarked, adding that many researchers in Chinese policy expect American companies will likely confront any genuinely novel frontier risk first.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe Chinese framework is confident they manage the application of these technologies within China,\u201d Webster continued. \u201cBeing online in China is linked to your real identity, and firms and users can be held liable.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Webster speculated that the same mechanisms model providers employ to refuse engagement on certain political subjects may be adapted to ensure models decline to execute offensive cyber operations or provide adverse biological engineering results. He noted that since Chinese firms typically coordinate with regulators behind closed doors, it can be challenging to ascertain what internal evaluations they perform prior to public release.<\/p>\n<p class=\"wp-block-paragraph\">Proponents of open-weight AI contend that disclosing the weights is vital for cybersecurity because it enables companies to protect themselves against assaults \u2014 Hugging Face utilized GLM-5.2 to defend against OpenAI\u2019s breach \u2014 and ensures they are better equipped to confront future threats by being aware of what is anticipated.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe identical systems that aided in thwarting an AI-driven cyberattack can now help to defend against countless cyber threats daily while assisting us in identifying and rectifying vulnerabilities before they can be exploited by attackers,\u201d Clem Delangue, CEO of Hugging Face, stated this week in a social media update.<\/p>\n<p class=\"wp-block-paragraph\">Papadatos remarked that this advantage is often exaggerated and doesn&#8217;t imply \u201cwe should make dangerous capabilities open source.\u201d <\/p>\n<p class=\"wp-block-paragraph\">\u201cThe critical point for me is that we must not accept that dangerous capabilities are readily accessible to anyone, anywhere,\u201d he emphasized, asserting that the industry should aim to restrict access to only \u201cgood capabilities.\u201d Attackers typically adapt to new tools at a faster rate than defenders. For instance, a ransomware group can modify its tactics within a week, while a hospital cannot.\u201d<\/p>\n<\/div>\n<p><em>When you purchase through links in our articles, we may earn a small commission. This doesn\u2019t affect our editorial independence.<\/em><\/p>\n","protected":false},"author":2,"featured_media":3491524,"comment_status":"open","ping_status":"closed","sticky":false,"template":"Default","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-3491523","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/techingeek.com\/index.php\/wp-json\/wp\/v2\/posts\/3491523"}],"collection":[{"href":"https:\/\/techingeek.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/techingeek.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/techingeek.com\/index.php\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/techingeek.com\/index.php\/wp-json\/wp\/v2\/comments?post=3491523"}],"version-history":[{"count":0,"href":"https:\/\/techingeek.com\/index.php\/wp-json\/wp\/v2\/posts\/3491523\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/techingeek.com\/index.php\/wp-json\/wp\/v2\/media\/3491524"}],"wp:attachment":[{"href":"https:\/\/techingeek.com\/index.php\/wp-json\/wp\/v2\/media?parent=3491523"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/techingeek.com\/index.php\/wp-json\/wp\/v2\/categories?post=3491523"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/techingeek.com\/index.php\/wp-json\/wp\/v2\/tags?post=3491523"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}