How AI safeguards are hindering the efforts of offensive cybersecurity researchers

How AI safeguards are hindering the efforts of offensive cybersecurity researchers

For several months, major AI companies have implemented specially approved programs and stringent safeguards to restrict the use of their models by malicious actors. However, these restrictions are now impairing the efforts of legitimate network defenders as well as offensive cybersecurity researchers. 

In June, the U.S. government imposed export control measures on Anthropic’s widely discussed AI models, Mythos and Fable. This action was partly triggered by a report suggesting that it might be feasible to circumvent the models’ safeguards designed to stop users from employing them to create and carry out harmful cyberattacks.

Regardless of whether the situation was genuinely influenced by concerns regarding a jailbreak, it’s evident that Anthropic has frequently promoted Mythos as a doomsday cyber tool that can only be entrusted to thoroughly vetted users, and even then with rigorous safeguards enforced. (The export restrictions on Fable 5 and Mythos 5 have since been rescinded. Fable 5 returned to public access on July 1; Mythos 5 has been reintegrated solely to screened U.S. organizations as part of the government’s evaluation process.)

This type of gatekeeping is not exclusive to Mythos. Both Anthropic with its other models and OpenAI offer cybersecurity researchers avenues to apply for vetting, and — if approved — gain access to models with reduced cybersecurity limitations: OpenAI’s Trusted Access for Cyber initiative and Anthropic’s Cyber Verification Program. 

These safeguards have faced significant criticism, especially from researchers tasked with uncovering unknown vulnerabilities within systems and formulating ways to exploit them before malicious actors do.

In a recent appearance on a cybersecurity podcast, Mark Dowd, a renowned security researcher, expressed that “it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not.”

Dowd has spent years identifying and selling “zero-days” — previously unidentified software flaws and the exploits that leverage them — to Western governments instead of reporting them to the software developers for patching. Governments pay a premium for vulnerabilities because they remain unaddressed, beneficial for intelligence activities.

Dowd acknowledged that his work might color his perspective, but he is not the only one. Multiple individuals working in offensive cybersecurity — who actively probe systems for weaknesses — described to TechCrunch their use of AI tools and how they navigate their guardrails. 

Chris Anley, the chief scientist at security consulting powerhouse NCC Group, mentioned that prompting an AI model to attempt to exploit a bug is crucial in verifying that it’s a genuine vulnerability and worth addressing. However, if a guardrail leads the model to outright refuse to respond, it adversely affects defenders, he stated.

“This is where the entire offensive versus defensive and guardrails aspect comes in, because ‘fix this code’ as a prompt serves as both a vital mechanism for defense and a blueprint for identifying critical vulnerabilities within the code base,” Anley explained. “Thus, the same tool functions as both an offensive and a defensive tool, and the two cannot actually be separated.”

“It’s ‘like a hammer,’” he continued. “You cannot construct a house without a hammer. It is undoubtedly a tool but irreducibly also a weapon.”

When he and his team encounter such hurdles, they sometimes revert to open-source AI models that come with no safeguards whatsoever.

Paolo Stagno, the chief technology officer at Crowdfense, a well-known entity that develops, procures, and sells undiscovered vulnerabilities to government bodies, shared Dowd’s sentiments, stating that AI companies “essentially treat clients like children who require supervision” with their vetted programs and guardrails. 

Stagno stated that he and his team do utilize frontier models — but strictly for reverse engineering. They refrain from employing AI to assist in identifying vulnerabilities or constructing exploits, as integrating that work into a cloud-based model poses risks of exposing sensitive vulnerability information or having it absorbed into future training sessions. For that purpose, he stated they resort to locally run open-source models, avoiding data sharing outside the model. 

Giuseppe Cali, a security researcher specializing in zero-days and exploit development, claimed that guardrails do not obstruct his work. This is because he does not use AI for offensive tasks; rather, he employs it for initial reverse engineering, to comprehend the code he’s analyzing, and to create supporting tools. For that, he emphasized that AI tools can expedite the process and allow him to concentrate on discovering vulnerabilities. 

“I still want to control the actual bug discovery and weaponization myself, and that wouldn’t change if all guardrails were removed tomorrow,” Cali remarked. “I am possessive about my bugs, and I enjoy this game too much to allow models to play it for me.”

One researcher at a smartphone-component manufacturer, who requested anonymity due to lack of authorization to speak to the media, indicated that his employer is not part of Anthropic’s CVP program, leaving its tools nearly ineffective for discovering vulnerabilities due to overly strict guardrails.

“If it catches wind we’re doing anything security related, it just stops and isn’t usable,” the person remarked. 

Chris Thompson — chief executive of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an event focused on offensive security and AI — stated that in his experience with frontier AI models, the guardrails can be erratic and vary in performance daily. This inconsistency holds true even within the more lenient frameworks of Anthropic’s and OpenAI’s vetted programs. 

“I think the practical effect is that you spend a lot of time negotiating with the model instead of concentrating on the fundamental security program,” Thompson remarked. “Rather than analyzing a vulnerability and reasoning through the exploitability, you are attempting to uncover why you are encountering inconsistent results or why models are over-sanitizing the output.” 

Consequently, researchers are driven towards or rely on Chinese open-source models like GLM — models that can be freely downloaded and executed locally without any vetting or usage constraints — noted Thompson.

“You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,” he remarked. “I believe it’s more damaging than beneficial to have these guardrails in place.”

Instead of imposing further restrictions, Thompson urged the AI frontier labs to expand their programs, provide responsible access, and hold those who misuse their tools accountable. Otherwise, he contended, defenders will lose the AI race.

“There’s this big storm approaching. There’s this significant wave of attacks that are going to occur at unprecedented speed and scale,” Thompson warned. “But the same security consulting companies and legitimate researchers striving to make a difference are currently being stifled.”

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

Leave a Reply