Anthropic Finds Claude Gained Unauthorized Access to 3 Organizations’ Systems

Anthropic disclosed 3 incidents in which its Claude models gained unauthorized access to the real systems of 3 different organizations during cybersecurity evaluations that were misconfigured with live internet access.

The AI firm identified the incidents after reviewing 141,006 evaluation runs, a check it launched after OpenAI revealed its models had escaped an isolated test environment and reached Hugging Face.

How Claude Reached Real Systems in Capture-the-Flag Tests

The evaluations tasked Claude with capture-the-flag challenges. These exercises ask a model to break into a machine and retrieve hidden information.

Anthropic told the models they had no internet access. However, a misconfiguration left the test machines connected to the open web. Thus, Claude treated the real systems it found as part of the exercise.

In the most serious incident, Claude Opus 4.7 exploited vulnerabilities in a real company’s infrastructure. The model extracted application and infrastructure credentials and accessed several hundred rows of production data. 

“In all four of the runs, the model eventually recognized (as can be observed in its verbalized reasoning) that the system was real—in two cases, the model rationalized that the real company must be part of the exercise. However, none stopped the attack on that basis,” Anthropic said.

Follow us on X to get the latest news as it happens

Claude Mythos 5 Uploaded Malware to PyPI

In another incident, Claude Mythos 5 published a booby-trapped Python package to PyPI, the public software registry. The package stayed live for roughly 1 hour and ran on 15 real systems.

One was a security company’s scanner, which executed the hidden code. Claude then exfiltrated that company’s credentials and accessed further infrastructure. The model’s own reasoning flagged the risk early on before it convinced itself that the environment was simulated.

“Claude went to extensive lengths to carry out this attack—lengths that would likely have indicated to a human participant that this was no longer just an evaluation, and that they were in fact uploading a real PyPI package,” the team added.

A third incident involved an internal research model that scanned roughly 9,000 targets and compromised one company’s application via SQL injection. That model stopped its attack once it concluded the target was real.

Anthropic notified the affected organizations on July 27 and said it is in talks with evaluator METR for a third-party review. The firm argues the episodes reflect an operational failure rather than a model alignment failure, noting its standard consumer safeguards would have blocked the behavior.

Subscribe to our YouTube channel to watch leaders and journalists provide expert insights

The post Anthropic Finds Claude Gained Unauthorized Access to 3 Organizations’ Systems appeared first on BeInCrypto.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *