AI model goes rogue after ‘escaping’ into the internet and launching cyberattack

OpenAI, the company behind ChatGPT, has revealed that one of its agents was able to ‘escape’ onto the internet during a recent data test, an event it has called an ‘unprecedented’ breach.

Confirmed in a press release shared with the public yesterday (21 July), the company explained that it had been running a controlled test with its latest GPT-5.6 Sol agent and a newer model yet to be released to the public when the models were able to gain access to the internet, after identifying vulnerabilities within the company’s research environment.

From here, the agents then accessed databases run by Hugging Face, a separate platform in search of answers to the test, therefore ‘cheating’ the system.

The breach was ultimately discovered when servers from both OpenAI and Hugging Face separately spotted ‘anomalous activity’.

OpenAI, the parent company of ChatGPT, confirmed the news yesterday (Getty Stock Images)

OpenAI, the parent company of ChatGPT, confirmed the news yesterday (Getty Stock Images)

“All evidence suggests that the models were hyper-focused on finding a solution for ExploitGym [the task it had been set], going to extreme lengths to achieve a rather narrow testing goal,” read a statement from the company.

“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly.”

So, in short, when artificial intelligence isn’t preparing to take over our jobs or being used to create explicit images of famous people, advanced models are hacking their way out of a system and gaining access to ‘secret information’ that it could use to cheat the evaluation.

OpenAI has added that it is now working with Hugging Face in order to obtain a greater understanding of what advanced AI models are capable of.

“We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret,” Clem Delangue, Co-founder and CEO, Hugging Face, said in a statement.

“It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”

This wouldn’t be the first time AI has gone rogue in recent months either, with a recent study published by The Guardian revealing that an increasing number of chatbots have been found to lie or cheat in the past six months.

How the latest AI breach sounds (Getty Stock Images)

How the latest AI breach sounds (Getty Stock Images)

However, Delangue has insisted he believes there was ‘no malicious intent’ on OpenAI’s part in a statement on X, writing: “We’ve spent the past 24 hours working closely with the @OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part.

“It’s quite mind-blowing that all of this happened autonomously!”

OpenAI has also released a list of actions which it is taking in order to improve AI controls. “As part of the investigation, we are implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched,” the statement read.

“We are regularly briefing our Safety and Security Committee on these controls and their impact.”

The company added that it is also working on creating ‘stronger protections’ in future training, saying: “This incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time, and monitoring during internal testing.”

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *