Tech

AI Guardrails Stifle Defensive Cybersecurity Efforts, Researchers Warn

Industry experts argue that arbitrary restrictions on AI tools are obstructing offensive security work, with some researchers citing inconsistent outputs and data leakage risks as they pivot to Chinese systems like GLM.

Author
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: TechCrunch · original
How AI guardrails are impeding the work of offensive cybersecurity researchers
Strict safety protocols from OpenAI and Anthropic are hindering legitimate vulnerability research, pushing defenders toward foreign open-source models.

Cybersecurity researchers are reporting that strict guardrails implemented by major artificial intelligence providers, including OpenAI and Anthropic, are impeding their ability to identify vulnerabilities and develop exploits. While these measures were designed to prevent malicious use, experts argue they are obstructing legitimate defensive efforts, forcing some professionals to rely on open-source models or foreign systems such as GLM. Critics warn that such restrictions may stifle the ability of defenders to keep pace with emerging cyber threats.

The U.S. government imposed export control restrictions on Anthropic’s AI models Mythos and Fable in June, citing concerns that the models’ guardrails could be bypassed to build malicious cyberattacks. These controls have since been adjusted, with Fable 5 returning to general access on July 1 and Mythos 5 reintroduced only to vetted U.S. organisations as part of a government review process. Despite these adjustments, gatekeeping remains a significant point of contention for security professionals who require unrestricted access to frontier models for vulnerability discovery.

Mark Dowd, a security researcher who sells zero-day vulnerabilities to Western governments, criticised AI companies for making arbitrary decisions about what constitutes safe security work. Dowd noted that it is uncomfortable for large corporations to determine what is safe in security, a sentiment echoed by Chris Anley, chief scientist at NCC Group. Anley compared AI tools to a hammer, stating they are irreducibly both a weapon and a defensive mechanism, arguing that the same prompt used to fix code is also a roadmap for finding critical vulnerabilities.

Paolo Stagno, CTO at CrowdFense, described AI companies as treating customers like children who need babysitting. He noted that his firm avoids using cloud-based AI for vulnerability discovery to prevent data leakage, preferring open-source models run locally. Conversely, Giuseppe Cali, a security researcher, stated that guardrails are not impeding his work because he uses AI only for initial reverse engineering and tool building, preferring to retain ownership of bug discovery and weaponisation.

Chris Thompson, CEO of RemoteThreat, described guardrails as inconsistent, noting that researchers spend significant time negotiating with the model rather than analysing vulnerabilities. Thompson warned that researchers are being pushed toward Chinese open-source models like GLM, which he argued is more harmful than beneficial. He called for AI frontier labs to open up their programs and provide responsible access, warning that defenders risk losing the AI race if legitimate researchers continue to be stifled.

Continue reading

More from Tech

Read next: TechCrunch and Stripe to identify Australia’s next breakout startups in Sydney
Read next: Open-source library 98.css brings Windows 98 aesthetic to modern web development
Read next: FDA Panel Narrowly Backs Reclassification of Controversial Peptides