● News

AI Guardrails Hampering Legitimate Cybersecurity Research Work

July 24, 2026 · snewle
AI Guardrails Hampering Legitimate Cybersecurity Research Work

AI guardrails designed to prevent malicious use are creating significant obstacles for offensive cybersecurity researchers and legitimate network defenders. While AI companies have implemented strict restrictions and vetted programs to limit model abuse, these measures are now hindering professionals whose job is to find unknown vulnerabilities before criminals can exploit them.

What Sparked the Guardrail Controversy?

In June, the U.S. government imposed export control restrictions on Anthropic’s AI models Mythos and Fable. The action followed a report demonstrating possible bypasses of guardrails designed to prevent users from building and executing malicious cyberattacks. Anthropic had marketed Mythos as requiring careful user vetting and strict guardrails. The export controls on Fable 5 and Mythos 5 were later lifted, with Fable 5 returning to general access on July 1, while Mythos 5 became available only to vetted U.S. organizations.

Both Anthropic and OpenAI now offer specialized programs for cybersecurity researchers: OpenAI’s Trusted Access for Cyber and Anthropic’s Cyber Verification Program. These programs provide approved users with access to models featuring fewer cybersecurity restrictions.

How Are Researchers Being Affected?

Security researcher Mark Dowd, known for finding and selling zero-days to Western governments, criticized the approach during a podcast appearance, stating it’s uncomfortable that “random large companies are making arbitrary decisions about what is safe in security and what’s not.”

Chris Anley, chief scientist at NCC Group, explained that asking AI models to exploit bugs is essential for confirming real vulnerabilities worth fixing. When guardrails block these queries, they harm defenders. He noted that prompts like “fix this code” serve as both defensive mechanisms and roadmaps for finding critical vulnerabilities, making the tools simultaneously offensive and defensive in nature.

Paolo Stagno, CTO at CrowdFense, said AI companies “essentially treat customers like children who need babysitting” with their vetted programs. His team uses frontier models only for reverse engineering, avoiding them for vulnerability discovery or exploit building due to concerns about leaking sensitive data through cloud-based models. Instead, they rely on open-source models run locally.

What Are the Practical Consequences?

Chris Thompson, CEO of RemoteThreat and founder of Offensive AI Con, reported that guardrails can be inconsistent and work differently daily, even within vetted programs. Researchers spend significant time “negotiating with the model instead of working on the core security program,” he said.

One researcher at a smartphone-component manufacturer, speaking anonymously, said his employer isn’t part of Anthropic’s CVP program, making the tools barely useful for finding vulnerabilities. “If it catches wind we’re doing anything security related, it just stops and isn’t usable,” the person stated.

Thompson warned that these restrictions push responsible researchers toward Chinese open-source models like GLM, which can be run locally without vetting or usage restrictions. “You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,” he explained, calling for AI labs to provide responsible access while holding abusers accountable.

Source: TechCrunch