OpenAI’s ‘rogue’ AI that hacked into Hugging Face did exactly what humans wanted it to — just not how they expected
When OpenAI recently revealed that two of its most advanced artificial intelligence (AI) models had escaped the confines of a cybersecurity test and hacked into a startup, it sounded a lot like the kind of scenario that AI safety researchers have spent years warning about.
The models found a previously unknown vulnerability in the infrastructure meant to contain them, gained access to the public internet and broke into Hugging Face, a major platform for hosting AI models and datasets. Their objective, however, was less sinister than the sequence of events might suggest: They were looking for information that would help them complete the cybersecurity test OpenAI had given them.
In a July 16 statement, Hugging Face representatives disclosed that internal datasets had been infiltrated, saying it was “different from anything we had handled before” because it was driven “by an autonomous AI agent system.” In another statement published July 21, OpenAI representatives fessed up to being responsible, calling the episode an “unprecedented cyber incident” while warning that similar events could become more common as AI models become increasingly capable.
Latest Videos FromLive Science
Independent experts who spoke with Live Science agree that what happened is significant — but they cautioned against interpreting it as an AI system suddenly developing a malicious agenda. The models appear to have pursued the task OpenAI gave them, finding a route to success that their creators had failed to anticipate or adequately block.
“If there’s a failure here, it isn’t that the AI wanted to hack something,” Oli Buckley, a professor in cybersecurity at Loughborough University in the U.K., told Live Science. “It’s that humans created a test where success was measured by achieving an objective, deliberately relaxed some of the normal security controls to measure the system’s capabilities, and underestimated how effective the model would be at finding an unexpected path to success.”
How did an OpenAI test end up hitting Hugging Face like this?
OpenAI was testing GPT-5.6 Sol and a more powerful unreleased model using ExploitGym, a benchmark that challenges AI systems to find and exploit software vulnerabilities. The company removed some cybersecurity safeguards that would normally prevent potentially dangerous actions while relying on an isolated environment to keep the models away from the wider internet.
According to OpenAI’s postmortem, the models discovered a previously unknown vulnerability in third-party software used to proxy and cache software packages. They exploited it, escalated their privileges and moved through OpenAI’s research infrastructure until they reached a machine with public internet access.
Hugging Face became a target because the models identified it as a possible source of information that could help them complete the ExploitGym challenges. OpenAI said at least one attack chain involved stolen credentials and previously unknown vulnerabilities that eventually enabled the models to execute remote code on Hugging Face systems and access test solutions stored in a production database.
In their disclosure, Hugging Face representatives said the company recorded more than 17,000 actions during the intrusion, but they couldn’t initially explain who or what was behind it. OpenAI’s subsequent disclosure supplied that missing piece: Its models had broken out of their test environment and gone looking for the answers elsewhere.
Rather than harboring any malicious intent, the AI models simply wanted to find out more information so they could complete their task.
(Image credit: wildpixel/ Getty Images)
Did the AI really “escape”?
It’s notable that the models found a flaw in the infrastructure designed to contain an AI and used it to reach the public internet. Describing the models as having “gone rogue,” however, risks assigning them unsupported motivations, Buckley said.
“I think I’d be wary of jumping to “rogue AI,”” Buckley said. “The models didn’t develop their own agenda or decide to attack Hugging Face while twirling their digital moustache.”
Buckley compared it to asking a dog to fetch a ball while leaving the garden gate open. “If the easiest ball for it to find is in the park down the road, that’s where it’ll head,” he said. “You wouldn’t say the dog had gone rogue; you’d just say you underestimated how literally it would pursue the task.”
Daniel Hulme, entrepreneur in residence at University College London and CEO of AI safety company Conscium, agreed that the models shouldn’t be assigned human-like motivations. “Models don’t have intent; humans have the intent, and we train models with goals in mind,” he told Live Science
The capability may matter more than the motive
What matters more than the models’ supposed motives is what they managed to accomplish while pursuing their assigned task.
“The genuinely significant point is that the models appear to have chained together multiple vulnerabilities across different systems and sustained a complex sequence of actions,” Buckley said. “That demonstrates a level of capability that security professionals should take seriously.”
The lesson isn’t that AI has become malicious. Instead, it’s that increasingly capable systems will exploit opportunities that humans fail to anticipate.
Oli Buckley, professor in cybersecurity at Loughborough University
Katerina Mitrokotsa, a professor of cybersecurity and applied cryptography at the University of St. Gallen in Switzerland, said the containment failure is particularly concerning because another company ultimately paid the price.
“What concerns me most is who ended up affected,” Mitrokotsa told Live Science. “The victim was not the company running the test, but a third party. This is the scenario security researchers have warned about for some time: that an AI agent’s escape does not necessarily stay contained to the environment in which it originated.”
OpenAI representatives said they have tightened the infrastructure used for these evaluations. But Mitrokotsa warned that containment becomes harder to guarantee as models improve at performing exactly the kind of exploitation OpenAI was testing.
An AI warning — and an impressive product demonstration
There is also reason to look carefully at how the incident is being framed. OpenAI’s account serves two purposes at once: It warns about the security risks posed by increasingly capable AI while demonstrating just how capable its own newest models have become.
Buckley said announcements from frontier AI companies like OpenAI or Anthropic should be viewed in the context of an industry competing to build ever-more-powerful models.
“We’ve seen similar high-profile capability demonstrations from Anthropic and others,” he said. “That doesn’t make the findings untrue, but it does mean we should separate the technical evidence from the marketing narrative.”
These companies have every incentive to show both that their models are extraordinarily capable and that they are taking the risks seriously, he added. The Hugging Face incident demonstrates both that OpenAI’s models carried out a complex series of operations with considerable autonomy and that its security measures failed to keep them inside the experiment.
Hulme argued that the longer-term challenge is ensuring that increasingly capable AI systems pursue their goals in ways that remain consistent with human values.
“Rather than seeking to control AIs, the focus should instead be on alignment,” he said, adding that continuous testing will be needed to ensure systems remain aligned with their intended missions while staying secure.
The episode, the experts said, leaves OpenAI with a result that is impressive and uncomfortable in equal measure. Its models found previously unknown vulnerabilities and continued pursuing their goal well beyond the boundaries their creators expected, but none of that requires them to have developed malign intentions.
“The lesson isn’t that AI has become malicious,” Buckley said. “Instead, it’s that increasingly capable systems will exploit opportunities that humans fail to anticipate.”
In this incident, OpenAI’s new models were given a hacking challenge and they were rewarded for finding a way to solve it. The humans running the experiment simply hadn’t anticipated quite how far they might go.