Close Menu
  • Home
  • AI
  • Entertainment
  • Finance
  • Sports
  • Tech
  • USA
  • World
  • Latest News

Subscribe to Updates

Subscribe to our newsletter and never miss our latest news

What's Hot

AMD takes on Nvidia with Helios AI rack-scale system

July 24, 2026

US, other countries back open source AI with ‘strong security’ at China summit

July 24, 2026

Vanderpump Rules’ Tom Schwartz slams paparazzi photos as ‘unrecognizable’

July 24, 2026
Facebook X (Twitter) Instagram
Facebook X (Twitter) Instagram Vimeo
BWE News – USA, World, Tech, AI, Finance, Sports & Entertainment Updates
  • Home
  • AI
  • Entertainment
  • Finance
  • Sports
  • Tech
  • USA
  • World
  • Latest News
BWE News – USA, World, Tech, AI, Finance, Sports & Entertainment Updates
Home » How AI guardrails are hindering the work of offensive cybersecurity researchers
AI

How AI guardrails are hindering the work of offensive cybersecurity researchers

adminBy adminJuly 24, 2026No Comments6 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr WhatsApp VKontakte Email
Share
Facebook Twitter LinkedIn Pinterest Email


For months, the AI ​​giant has devised special, vetted programs and strict guardrails to limit the use of its models by malicious hackers. However, these limitations currently impede the work of offensive cybersecurity researchers as well as legitimate network defenders.

In June, the US government imposed export control restrictions on Anthropic’s highly touted AI models Mythos and Fable. The move was prompted, at least in part, by a report that claimed it was possible to bypass model guardrails designed to prevent users from using the model to construct and execute malicious cyberattacks.

Regardless of whether this incident was truly motivated by fear of jailbreak, the fact is that Anthropic has repeatedly promoted Mythos as some sort of apocalyptic cybermachine that can only be made available to carefully vetted users, and with strict guardrails in place. (Export restrictions on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to public access on July 1. Mythos 5 was only reintroduced to vetted U.S. organizations as part of a government review process.)

This kind of gatekeeping is not unique to Mythos. Anthropic and Other Models and OpenAI both offer programs that give cybersecurity researchers access to less restrictive cybersecurity models if vetted and approved: OpenAI’s Trusted Access for Cyber ​​and Anthropic’s Cyber ​​Verification Program.

These guardrails have been widely criticized, especially by researchers whose job is to discover unknown vulnerabilities in systems and devise ways to exploit them before criminals can attack them.

Mark Dowd, a prominent security researcher, said in a recent appearance on a cybersecurity podcast that “I’m not very comfortable with these random big companies making arbitrary decisions about what’s security-wise and what’s not.”

For decades, Dowd has been discovering “zero days” — previously unknown software flaws and exploits — and selling them to Western governments rather than reporting them to software manufacturers to patch them. Governments pay a premium for vulnerabilities because vulnerabilities that serve intelligence operations remain open.

Dowd acknowledged that his job can create bias, but he’s not alone. Several people involved in offensive cybersecurity, who actively probe systems for weaknesses, explained to TechCrunch how they use AI tools and address guardrails.

Chris Unley, principal scientist at security consulting giant NCC Group, said attempting to exploit bugs in AI models is an important step to confirming that they are real vulnerabilities worth fixing. But guardrails can hurt defenders if they cause the model to refuse to fully answer questions, he said.

“This is where the whole attack and defense and guardrails part becomes important, because the prompt, ‘Fix this code,’ is not only an essential mechanism for defense, but it is also a roadmap for discovering critical vulnerabilities in your code base,” Anley said. “So the same tool is both an offensive tool and a defensive tool, and you can’t really choose between the two.”

It’s “like a hammer,” he continued. “You can’t build a house without a hammer. A hammer is definitely a tool, but it’s also a weapon.”

When he and his colleagues encounter such obstacles, they often turn to open-source AI models that have no guardrails.

Paolo Stagno, chief technology officer at Cloudfence, a well-known company that develops, acquires and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying that with vetted programs and guardrails, AI companies are “basically treating their customers like children who need babysitting.”

Stagno said he and his colleagues do use the Frontier model, but only for reverse engineering. He said they avoid using AI to find vulnerabilities or build exploits. Inputting that work into a cloud-based model risks exposing sensitive vulnerability data or absorbing it into future training runs. He said that step uses an open source model that runs locally because it doesn’t rely on data sharing outside the model.

Giuseppe Cali, a security researcher who discovers zero-days and develops exploits, said the guardrails have not hindered his work. That’s because he doesn’t use AI for offensive work. Instead, we use it for initial reverse engineering, understanding the code we’re analyzing, and building supporting tools. To that end, he said, AI tools speed up the process and allow them to focus on finding vulnerabilities.

“I still want to do the actual bug discovery and weaponization myself, and even if all the guardrails were lifted tomorrow, that wouldn’t change,” Cali said. “I’m jealous of my bugs, but I love this game too much to have a model play it.”

A researcher at a smartphone parts maker, speaking on condition of anonymity because he was not authorized to speak to the press, said his employer is not part of Anthropic’s CVP program, so the guardrails are too strict and the tool is of little use in finding vulnerabilities.

“When the wind blows, we do security-related things, and the wind stops and we can’t use it,” the official said.

Chris Thompson, CEO of cybersecurity company RemoteThreat and founder of Offensive AI Con, an event focused on offensive security and AI, said that from his experience with frontier AI models, guardrails are inconsistent and can behave differently every day. This is true even within the looser boundaries of Anthropic and OpenAI’s vetted programs.

“I think the practical impact is that you spend a lot of time negotiating with models instead of working on your core security program,” Thompson says. “Rather than analyzing vulnerabilities and reasoning through exploitability, we’re trying to find out why we’re getting inconsistent results or why the model over-sanitizes the output.”

As a result, researchers rely on or are pushed by Chinese open source models like GLM, which are freely downloadable models that can be run locally without scrutiny or usage restrictions, Thompson said.

“Responsible researchers are being forced out of U.S. government systems and into foreign-owned systems,” he said. “I think putting up these guardrails will do more harm than good.”

Thompson called on the AI ​​Frontier Institute to make its programs public, provide responsible access, and hold those who abuse its tools accountable, rather than further tightening regulations. Otherwise, he argued, defenders will lose the AI ​​race.

“There’s a big storm coming. There’s going to be a big wave of attacks at a speed and scale that we’ve never seen before,” Thompson said. “But those same security consulting firms and legitimate researchers who are trying to make a difference are now being suppressed.”

If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.



Source link

Share. Facebook Twitter Pinterest LinkedIn Tumblr WhatsApp Email
Previous ArticleAI ETF gains big despite tough quarter
Next Article Japan’s core inflation rate gradually rose from four-year low in June due to soaring oil prices
admin
  • Website

Related Posts

AMD takes on Nvidia with Helios AI rack-scale system

July 24, 2026

Meta launches new AI optimism ad set to song about human extinction

July 24, 2026

OpenAI makes ChatGPT Health available to all users in the US

July 23, 2026

As generated media becomes crowded, Runway launches AI model router

July 23, 2026
Leave A Reply Cancel Reply

Our Picks

Newly freed hostages face long road to recovery after two years in captivity

October 15, 2025

Former Kenyan Prime Minister Raila Odinga dies at 80

October 15, 2025

New NATO member offers to buy more US weapons to Ukraine as Western aid dwindles

October 15, 2025

Russia expands drone targeting on Ukraine’s rail network

October 15, 2025
Don't Miss
Entertainment

Vanderpump Rules’ Tom Schwartz slams paparazzi photos as ‘unrecognizable’

By adminJuly 24, 20260

Tom Schwartz knows it’s time to clear the air. When a tabloid published a paparazzi…

Behind the scenes of Chadwick Boseman and Taylor Simone Ledward’s private love story

July 24, 2026

Bachelorette producer Julie LaPlaca talks all about Peter Weber’s romance

July 24, 2026

Jax Taylor dates Brittany Cartwright’s publicist: Timeline revealed

July 23, 2026
About Us
About Us

Welcome to BWE News – your trusted source for timely, reliable, and insightful news from around the globe.

At BWE News, we believe in keeping our readers informed with facts that matter. Our mission is to deliver clear, unbiased, and up-to-date news so you can stay ahead in an ever-changing world.

Our Picks

Is it hot in the city? Free movies help Italians ‘feel at home’ in Rome

July 23, 2026

Saudi-US nuclear deal raises concerns about Middle East arms race

July 23, 2026

‘Caught in smoke’: Attack on ‘Russian Amazon’ brings Putin’s war closer to home

July 23, 2026

Subscribe to Updates

Subscribe to our newsletter and never miss our latest news

Facebook X (Twitter) Instagram Pinterest
  • Home
  • About Us
  • Advertise With Us
  • Contact US
  • DMCA
  • Privacy Policy
  • Terms & Conditions
© 2026 bwenews. Designed by bwenews.

Type above and press Enter to search. Press Esc to cancel.