An advanced artificial intelligence model built by ChatGPT maker OpenAI broke out of a secure test environment and attacked Hugging Face, a New York-based AI startup, OpenAI has admitted.
Hugging Face co-founder Thomas Wolf said the incident should serve as "a wake-up call" for the wider tech industry.

Wolf told BBC's Newsday radio programme that AI-driven attacks will soon become "one of the most common types of cyber-attacks we see." He said most companies are currently unprepared for the threat and do not realise that "the game has changed."
OpenAI described the episode as an "unprecedented cyber incident" involving state-of-the-art cyber capabilities. The company said its agent became so fixated on cheating a cybersecurity test that it broke out of its secure sandbox to steal answers from Hugging Face's systems. According to AI experts, the entire episode unfolded without any human intervention.

How the AI broke containment
OpenAI said the intrusion was caused by a combination of its models, including its newly released GPT-5.6 Sol and an "even more capable" model still being tested internally. The models had been set a standard cybersecurity benchmark test designed to evaluate their hacking abilities.
Rather than solving the tasks directly, the AI became "hyperfocussed" on cheating the test by accessing the internet. The agent first hacked OpenAI's own systems, moving from computer to computer until it found a "node" with internet access.

Hugging Face, one of the largest online platforms for sharing open-source AI models and a key resource for developers and researchers, became a prime target for the bot's search for answers.
Hugging Face's response
Wolf said the company initially had no idea where the attack was coming from when signs of disturbance emerged in mid-July. AI experts at the company said the attack was "very different" to anything the site had witnessed before.

Wolf said there were 17,000 attacks on Hugging Face's network, all coming from different IP addresses in a "very short time." OpenAI eventually realised what was happening and informed Hugging Face that its model was behind the attack, but not before the AI used stolen credentials and discovered a previously unknown vulnerability to access the startup's servers.
The UK's AI Security Institute is now studying how the AI system behaved during the incident and is working with OpenAI and other labs to strengthen safeguards, according to the BBC. The attack was especially concerning because the AI appears to have deliberately ignored or avoided usual safeguards in pursuit of what was a fairly routine task.

Experts sound alarm
Andrea Miotti, founder and chief executive of AI risk non-profit ControlAI, told the Daily Mail: "We can expect more of these rogue, fully autonomous AI attacks as AI companies continue trying to develop superintelligent AI, which is AI that could overpower our national security apparatuses and permanently evade human control."
Miotti added: "AI companies fundamentally do not understand how today's AIs work, and they have no idea how to control superintelligent systems vastly smarter than humans. Governments need to get pragmatic about this unprecedented risk and champion an international prohibition on developing superintelligence, before it's too late."
Richard Ford, chief technology officer at cybersecurity firm Integrity360, told the Daily Mail: "This is the moment many in cyber security have been warning about. Until now, we've seen attackers use AI to automate parts of an attack, but this is one of the first public examples of an AI agent independently identifying a weakness, escaping what should have been a secure environment and attempting to compromise another organisation."
OpenAI chief executive Sam Altman confirmed there had been a "significant security incident."
Second AI sandbox breach this year
The incident comes just months after OpenAI rival Anthropic revealed that its Mythos AI had broken out of its own safe "sandbox." Anthropic said the model had found thousands of high-severity vulnerabilities, including some in every major operating system and web browser.
Anthropic also described what it called "reckless destructive actions" by the model. The bot attempted to break out of its testing sandbox, hid its actions from researchers, broke into files that had been "intentionally chosen not to be made available," and posted exploit details publicly.
OpenAI has been contacted for comment.


