For months, leading artificial intelligence developers have implemented rigorous vetting processes and stringent safety protocols for their advanced models, ostensibly to prevent malicious actors from weaponizing these powerful tools. However, a growing chorus of cybersecurity professionals and researchers contends that these very safeguards are now inadvertently hindering the vital work of legitimate network defenders and offensive cybersecurity experts, potentially leaving critical systems more vulnerable. The intricate balance between controlling AI’s potential for harm and enabling its use for good has become a significant point of contention in the rapidly evolving landscape of digital security.
The Anthropic Export Control Incident and its Ripple Effects
A pivotal moment that brought this tension into sharp focus occurred in June when the U.S. government imposed export control restrictions on Anthropic’s highly anticipated AI models, Mythos and Fable. This decisive action was reportedly triggered, at least in part, by a report alleging that the models’ built-in guardrails, designed to prevent their use in constructing and executing cyberattacks, could be circumvented. While the precise motivations behind the government’s intervention remain a subject of debate, the practical outcome was a significant disruption for those seeking to utilize these powerful AI capabilities for cybersecurity research and defense.
Anthropic had consistently positioned Mythos as a potentially potent "doomsday cybermachine," emphasizing its release would be restricted to carefully vetted users with stringent safety measures in place. This marketing strategy, while intended to convey responsibility, also contributed to an atmosphere where the model was perceived as inherently dangerous, necessitating tight control. Although the export controls on Fable 5 and Mythos 5 have since been lifted, with Fable 5 returning to general access on July 1 and Mythos 5 reintroduced to vetted U.S. organizations following a government review, the incident underscored the challenges of managing access to powerful AI technologies.
The Rise of Vetted Access Programs and Their Criticisms
This form of gatekeeping is not exclusive to Anthropic’s Mythos. Both Anthropic and OpenAI have established dedicated programs designed to provide cybersecurity researchers with access to their AI models, albeit with fewer restrictions than typically found in general releases. OpenAI offers its "Trusted Access for Cyber program," while Anthropic provides a "Cyber Verification Program." These initiatives aim to equip ethical hackers and security professionals with advanced AI tools to discover and address vulnerabilities before they can be exploited by adversaries.
However, these meticulously crafted guardrails have become a significant point of friction for many in the cybersecurity community. Researchers whose primary role involves identifying previously unknown flaws in software systems and developing methods to exploit them – a crucial step in enabling timely patching – find these restrictions to be a considerable impediment.
Mark Dowd, a highly respected security researcher with decades of experience in discovering and selling "zero-day" vulnerabilities to Western governments, voiced his concerns during a recent cybersecurity podcast appearance. He articulated a sentiment shared by many: "it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not." Dowd’s work, which involves providing governments with intelligence on exploitable flaws that remain unpatched, highlights a different facet of cybersecurity where the existence of vulnerabilities is strategically valuable for intelligence operations, a practice that directly clashes with the AI developers’ goal of immediate remediation.
The Dual Nature of AI in Cybersecurity
The debate over AI guardrails is fundamentally rooted in the dual-use nature of these technologies in the cybersecurity domain. As Chris Anley, chief scientist at the prominent security consulting firm NCC Group, explained, the ability of an AI model to attempt to exploit a discovered bug is a critical step in validating its severity and the necessity of a fix. However, if an AI’s guardrails preemptively refuse to engage with such a request, they can inadvertently hinder the very defense mechanisms they are meant to protect.
"This is where the whole offensive versus defensive and guardrails part comes in, because ‘fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base," Anley stated. He further elaborated on this complex interplay, likening AI to a hammer: "You can’t build a house without a hammer. It’s definitely a tool but it’s also irreducibly a weapon as well." This analogy underscores the inherent challenge of separating the beneficial from the potentially harmful uses of AI in cybersecurity.
When faced with these restrictive guardrails, Anley and his colleagues have increasingly resorted to open-source AI models that offer no such limitations, allowing for more unfettered exploration of potential weaknesses.
Concerns Over "Babysitting" and Data Privacy
Paolo Stagno, chief technology officer at Crowdfense, a company specializing in the acquisition and sale of undisclosed vulnerabilities to government agencies, echoed Dowd’s sentiment. He characterized the AI companies’ approach, with their vetted programs and stringent guardrails, as essentially treating their customers "like children who need babysitting."
Stagno further elaborated on the practical implications for his work. While his team does utilize advanced AI models, their application is confined to reverse engineering tasks. They actively avoid using AI to assist in vulnerability discovery or exploit development. The primary concern is the risk of leaking sensitive vulnerability data or having it inadvertently absorbed into future AI training datasets when feeding such work into cloud-based models. For these critical, sensitive operations, Crowdfense relies on open-source AI models run locally, ensuring that data never leaves their controlled environment.
The Human Element in Vulnerability Discovery
Not all cybersecurity researchers find the current guardrails to be an insurmountable obstacle. Giuseppe Cali, a security researcher focused on zero-day discovery and exploit development, explained that while guardrails don’t impede his work, it’s because he strategically employs AI. He uses it for initial reverse engineering and to develop supporting tools, thereby accelerating the understanding of complex code. This allows him to dedicate his efforts to the core task of discovering novel vulnerabilities.
"I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow," Cali asserted. "I am jealous of my bugs, and I like this game too much to let models play it for me." This perspective highlights a segment of the cybersecurity community that values the intellectual challenge and proprietary nature of their discoveries, viewing AI as a supportive tool rather than a replacement for human ingenuity.
The Inconsistency of Guardrails and the Shift to Foreign Models
However, the experience of many remains one of frustration. A researcher at a smartphone-component manufacturer, who requested anonymity due to authorization restrictions, shared that their employer’s lack of participation in Anthropic’s Cyber Verification Program renders the AI tools "barely useful" due to overly strict guardrails. "If it catches wind we’re doing anything security related, it just stops and isn’t usable," the researcher stated.
Chris Thompson, CEO of cybersecurity firm RemoteThreat and founder of Offensive AI Con, a conference dedicated to offensive security and AI, has observed firsthand the inconsistency of these guardrails. He noted that even within the more permissive environments of Anthropic’s and OpenAI’s vetted programs, the behavior of AI models can vary significantly from day to day.
"I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program," Thompson explained. "Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output."
This unpredictability and the perceived over-sanitization of output are pushing researchers towards open-source AI models, particularly those originating from China, such as GLM. These models are freely downloadable, can be run locally without any vetting process, and come with no usage restrictions.
A Call for Openness and Accountability
Thompson expressed deep concern over this trend: "You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems." He argued that the current restrictive guardrails are ultimately "more harmful than good."
Instead of further tightening restrictions, Thompson advocates for a paradigm shift where AI frontier labs adopt a more open approach. He calls for enhanced programs that offer responsible access, coupled with robust mechanisms to hold those who abuse these powerful tools accountable. He warns that without such changes, the cybersecurity community risks losing the critical AI race.
"There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before," Thompson concluded. "But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now." The future of cybersecurity defense may well depend on how effectively AI developers can navigate the delicate balance between safety and enablement.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
