OpenAI, Anthropic Reveal Their Models Breached Organizations
Systems are Acting Autonomously and Breaching Test Environments
The OpenAI breach of Hugging Face unfolded like a horror movie. Hugging Face revealed it had been hacked in mid-July, and nearly a week later OpenAI stepped out of the shadows, and said its own systems had been responsible after breaking out of a testing environment. The combination of models had escaped by exploiting a previously unknown zero-day vulnerability in Hugging Face’s systems.
In a blog post, Hugging Face detailed the incident: “Over roughly two and a half days inside our infrastructure, an autonomous AI agent driven by a combination of OpenAI models ran an end-to-end intrusion against our platform: it was thousands of small, automated decisions, executed at machine speed across short-lived sandbox environments, with command-and-control staged on ordinary public web services.”
Following the incident, Anthropic said it conducted its own internal assessment on its own AI, Claude. “In particular, we looked for evidence that Claude – like the OpenAI models that accessed Hugging Face – was able to access the internet from within testing environments that should have been sealed off,” it wrote.
There, the company discovered several Claude models had breached three companies during security tests. The models accessed the internet from within or while interacting with third party, Irregular, a security partner of Anthropic. From there, the models gained unauthorized access to the production infrastructure of several companies.
Despite drawing on the OpenAI/Hugging Face incident when reporting on Claude, Anthropic clearly defined the differences between the incidents. While OpenAI’s models had exploited an unknown software vulnerability to break out of its test environment, Anthropic’s models had reached the internet through a connection left open by mistake.
Likewise, Anthropic noted it had discovered the incidents itself through review, whereas OpenAI had admitted to the breach following Hugging Face’s security announcement.