OpenAI agents breach Modal client system after Hugging Face hack

OpenAI’s ‘rogue agent’ took advantage of a code vulnerability, experts explained.

US cloud company Modal has confirmed that OpenAI’s agents were able to hack into one of its customer’s systems when the AI models breached containment and gained unauthorised access to Hugging Face earlier this month.

Last week’s incident sent shockwaves across the industry, raising serious concerns around AI’s rapidly advancing ability to bypass boundaries and, effectively, go “rogue”.

It comes amid increased scrutiny around OpenAI and Anthropic’s new AI models, resulting in gated launches and greater government involvement. Both AI giants have ramped up efforts to go public in blockbuster listings as they compete to gain market dominance and enterprise footing.

OpenAI CEO Sam Altman, in a recent interview, said that the Hugging Face breach was the first security incident he felt “viscerally”.

“I feel a little surprised that more people don’t feel it so viscerally,” he told Invest Like The Beast in a podcast episode published on Tuesday (28 July).

Hugging Face said that OpenAI’s agents accessed a sandbox “hosted on a ​third-party provider’s infrastructure” when it breached containment last week. A sandbox is an isolated environment where AI models are tested without production classifiers, or guardrails.

Modal chief technology officer Akshat Bubna confirmed that its customer set up a publicly accessible interface which enables anyone to use their sandbox.

“We’re aware a Modal customer published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution,” Bubna told Axios. “Their code had a vulnerability that was exploited…This was used by the rogue agent.”

“Modal’s platform was not compromised in any way,” he clarified.

In an updated statement, OpenAI said that none of its upcoming models were involved in exploiting Hugging Face. It explained that its models were able to identify and exploit a unknown zero-day vulnerability to gain access to the internet, which enabled it to access Hugging Face.

“In our ongoing review of the Hugging Face intrusion and broader activity from our models, we have been finding a small number of cases where the models identified and used publicly exposed credentials at the account-level on other publicly-available services.

“Based on our review to date, we have not identified any other activity at the level of severity or scale of what we’ve shared related to Hugging Face,” the company said.

Cybersecurity experts, however, believe that the breach is a result of “missing governance and control”.

“When conducting security testing you should define what is in and out of the testing scope, even for broad red team engagements,” said Richard Davies, director of cyber solutions, Talion.

“The reported impacts and timelines indicate this was not in place.”

CybaVerse chief technology officer Simon Phillips added: “The model, tooling and instructions were very loose, almost to the point it was told it could do anything on any system, which it clearly did.”

Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *