OpenAI says AI model went rogue, hacked Hugging Face
OpenAI on Tuesday said that its advanced artificial intelligence models went rogue and hacked into Hugging Face, a digital repository of AI technology.
"We had a significant security incident during evaluation of our models," OpenAI CEO Sam Altman said in a statement posted on social media.
The disclosure comes amid heightened concerns about the cybersecurity risks posed by powerful AI models.
In a blog post, OpenAI said the "unprecedented cyber incident" took place while it was testing the cybersecurity capabilities of its models in a "tightly controlled digital testing ground," where internet access was limited for safety.
It said the AI agent managed to escape containment, reach the internet and break into Hugging Face to try to satisfy its testing goal.
"While operating in our sandboxed testing environment, our models spent a substantial amount of (computing power) finding a way to obtain open Internet access, in pursuit of solving the evaluation problem," an OpenAI blog post about the incident said.
After connecting to the internet, the models decided to target the platform Hugging Face, a large repository of AI models, datasets and other information, to help in their quest.
AI under pressure: scams, security and sustainability
To view this video please enable JavaScript, and consider upgrading to a web browser that supports HTML5 video
AI agents are autonomous software systems that carry out tasks in the real world to achieve specific goals.
The AI firm said the incident involved a combination of models, including its recently launched GPT-5.6 Sol "and an even more capable pre-release model."
Hugging Face, an open-source AI platform, caused a stir in the cybersecurity community when it said in a blog post last week that it "was different from anything we had handled before" in that "it was driven, end to end, by an autonomous AI agent system."
In a post to X, Hugging Face co-founder Clement Delangue said the company suspected the hack "might have come from a frontier lab, given the sophistication of the agent. Turns out it did!"
He added: "It's quite mind-blowing that all of this happened autonomously!"
The disclosure comes amid growing concerns about the cybersecurity risks of advanced AI. In June, President Donald Trump ordered federal reviews of the most powerful AI systems before their public release.
OpenAI's most advanced AI model, GPT-5.6, was launched earlier this month, but only after its debut was delayed at the US government's request over national security fears.
In June, OpenAI rival Anthropic was forced to pull Fable 5 and cybersecurity Mythos model, over concerns that it could pose a security risk by giving hackers and other bad actors access to AI capable of finding and exploiting vulnerabilities exceptionally quickly.
Don't let the algorithm hide the news. If you rely on our team for trusted reporting, please take a moment to select us as your Preferred Source on Google by clicking here and hitting the "star" or "preferred" button, so you'll always see our verified news first.
Related Stories
AI News
This Is Probably Not the AI
25 seconds ago
AI News
Frontier AI labs still won’t say how they’d contain a rogue model
25 seconds ago
AI News
Ex
28 minutes ago
AI News
Editorial | As Hong Kong boosts governance efficiency with AI, balance is key
1 hour ago
AI News
Nvidia customers notified of AI
3 hours ago
AI News
Reality Check: A Guide to Navigating the AI Information Flood
4 hours ago
AI News
Protecting Digital Privacy In The Artificial Intelligence Era
4 hours ago
AI News
A third of web pages published since ChatGPT launched were written by AI, study finds
5 hours ago