AI Guardrails Stifle Defensive Cybersecurity Efforts, Researchers Warn
Industry experts argue that arbitrary restrictions on AI tools are obstructing offensive security work, with some researchers citing inconsistent outputs and data leakage risks as they pivot to Chinese systems like GLM.

Cybersecurity researchers are reporting that strict guardrails implemented by major artificial intelligence providers, including OpenAI and Anthropic, are impeding their ability to identify vulnerabilities and develop exploits. While these measures were designed to prevent malicious use, experts argue they are obstructing legitimate defensive efforts, forcing some professionals to rely on open-source models or foreign systems such as GLM. Critics warn that such restrictions may stifle the ability of defenders to keep pace with emerging cyber threats.
The U.S. government imposed export control restrictions on Anthropic’s AI models Mythos and Fable in June, citing concerns that the models’ guardrails could be bypassed to build malicious cyberattacks. These controls have since been adjusted, with Fable 5 returning to general access on July 1 and Mythos 5 reintroduced only to vetted U.S. organisations as part of a government review process. Despite these adjustments, gatekeeping remains a significant point of contention for security professionals who require unrestricted access to frontier models for vulnerability discovery.
Mark Dowd, a security researcher who sells zero-day vulnerabilities to Western governments, criticised AI companies for making arbitrary decisions about what constitutes safe security work. Dowd noted that it is uncomfortable for large corporations to determine what is safe in security, a sentiment echoed by Chris Anley, chief scientist at NCC Group. Anley compared AI tools to a hammer, stating they are irreducibly both a weapon and a defensive mechanism, arguing that the same prompt used to fix code is also a roadmap for finding critical vulnerabilities.
Paolo Stagno, CTO at CrowdFense, described AI companies as treating customers like children who need babysitting. He noted that his firm avoids using cloud-based AI for vulnerability discovery to prevent data leakage, preferring open-source models run locally. Conversely, Giuseppe Cali, a security researcher, stated that guardrails are not impeding his work because he uses AI only for initial reverse engineering and tool building, preferring to retain ownership of bug discovery and weaponisation.
Chris Thompson, CEO of RemoteThreat, described guardrails as inconsistent, noting that researchers spend significant time negotiating with the model rather than analysing vulnerabilities. Thompson warned that researchers are being pushed toward Chinese open-source models like GLM, which he argued is more harmful than beneficial. He called for AI frontier labs to open up their programs and provide responsible access, warning that defenders risk losing the AI race if legitimate researchers continue to be stifled.

