After OpenAI incident, Anthropic finds Claude hacked organisations

Anthropic said Claude was mistakenly given access to the internet.

Anthropic on Thursday (30 July) said it had found three instances where Claude models gained unintended access to the internet during cybersecurity evaluations prompted by a “misunderstanding” between the company and its testing partner Irregular.

The AI company said it launched a retrospective analysis of its testing systems on 23 July after rival OpenAI’s models were found to have hacked Hugging Face during testing earlier this month.

That breach had downstream consequences, when, earlier this week, US cloud company Modal revealed that the models also gained access to one of its customers.

In its analysis of more than 140,000 evaluation runs, Anthropic said it discovered three instances involving Opus 4.7, Mythos 5 and an internal research test model where the models broke through to the internet.

These occurred when the models were inside Irregular’s testing environment or interacting with it, Anthropic explained. The earliest incidents date back to April.

In one serious case, Opus 4.7 targeted a real company that shared names with a fictional company provided to it during testing, Anthropic said. Claude was able to extract application and infrastructure credentials from the business, and gained access to a database containing several hundred rows of production data, it added.

Anthropic explained that its test evaluation prompts explicitly did not allow internet access, but did not limit Claude’s reach. However, a misunderstanding between the company and Irregular left the machines conducting the tests with live internet. Neither party was aware of the errors until Anthropic’s analysis earlier this week, it said.

The Claude maker said it paused all cyber evaluations after identifying the breach and notified the three organisations its models hacked on Monday (27 July).

“Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone,” Anthropic wrote in yesterday’s blogpost.

Recent unintended cyberattacks carried out by powerful, ‘rogue’ agents have sent shockwaves across the AI industry, raising serious concerns around careful testing and models’ rapidly advancing ability to bypass boundaries.

“For threat actors with money to spend on tokens and access to less restricted models, the time taken to compromise a given target has likely reduced,” said Richard Davies, director of cyber solutions at Talion, earlier this week.

Hugging Face said that OpenAI’s agents accessed a sandbox hosted on a ​third-party provider’s infrastructure when they breached containment earlier this month. OpenAI maintained, in an updated statement, that none of its upcoming models were involved in the exploit.

Following the Hugging Face incident, members of the US Congress introduced a new bill which would require AI companies to be able to shut down, throttle or suspend their models if they go ‘rogue’.

However, some cybersecurity experts have said that missing governance and control is the reason behind the Hugging Face breach.

“The model, tooling and instructions were very loose, almost to the point it was told it could do anything on any system, which it clearly did,” said CybaVerse chief technology officer Simon Phillips.

“The story here isn’t about an AI model going rogue; the model did exactly what it was tasked to do.”

Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.

Dario Amodei at the World Economic Forum Annual Meeting. Image: 2026 World Economic Forum via Flickr (CC BY-NC-SA 4.0)

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *