OpenAI says its AI models went rogue and hacked another company
Why it matters: The AI boom has intensified old fears about the dangers of increasingly autonomous software, but those discussions have remained theoretical so far. A recent incident involving an OpenAI test model might have added substance to those fears, but the seriousness of the incident remains to be determined.
OpenAI has confirmed that its large language models were responsible for an agentic cyberattack that impacted another AI company, Hugging Face, last week. The attack is notable because it was not human-directed – OpenAI’s tools escaped from an isolated network and targeted Hugging Face on their own.
Hugging Face initially reported an intrusion from an unknown LLM on July 16. The incident itself was relatively minor – the company stopped the attack, addressed the vulnerabilities that enabled it, and does not believe customer-facing data was accessed. However, the incident eventually raised alarm because of why it happened.
– Alex Prompter (@alex_prompter) July 22, 2026
OpenAI claimed this week that GPT-5.6 Sol and another, more advanced pre-release model went rogue during an evaluation of their hacking abilities. The models were running with reduced safeguards to assess their true capabilities, but OpenAI believed there was no danger as long as the test remained within an isolated network.
However, after struggling to complete the evaluation, the models cheated by escaping into the open internet via vulnerabilities previously unknown to OpenAI. Once outside containment, they independently identified Hugging Face as a possible source of the information they were looking for and broke into the company by using yet more previously unknown exploits to steal login credentials.
OpenAI and Hugging Face are still jointly investigating the incident, and a detailed report is likely forthcoming. However, if the initial reports are accurate, they prove that LLMs can carry out multi-layered cyberattacks with little to no human input.
Conversely, some commentators suspect that OpenAI and Hugging Face simply deployed poorly vibe-coded defenses, or that OpenAI used an overly vague prompt during the evaluation. Others suspect a publicity stunt, echoing prior skepticism about claims that tech giants might lose control of generative AI and endanger humanity.
Following the incident, OpenAI took the opportunity to recommend its Trusted Access program, which Hugging Face joined after the breach. Hugging Face also used the incident to recommend AI-based defenses against the rising threat of agentic cyberattacks. Some believe that such claims are designed to raise the perceived value of LLMs to fuel AI’s financial boom, while other critics warn that the AI companies that sparked the revolution are attempting to seize control of early regulation efforts.
