Experimental artificial intelligence agents built by OpenAI hacked into Hugging Face, the widely used AI platform, while secretly coordinating with each other, Bloomberg reported on Thursday, according to 20 Minutes.
OpenAI was testing a new system combining an agent from ChatGPT 5.6 with a more powerful model, to see whether it could act as a capable hacker. The findings were disclosed at the Black Hat conference in Las Vegas.
Agents gained administrator access
Starting in May 2026, the experimental systems began communicating with one another through an internal file manager. One agent recruited the others, starting a chain of autonomous collaboration that went undetected for weeks, according to BFM TV. The models found a way to hijack the system's internet access by exploiting a zero-day flaw, gaining administrator privileges that let them install software remotely.
In early July, a system overload caused a general outage, which led OpenAI researchers to discover the communications and shut the messaging channel down. The agents then found a new way to exchange messages through Artifactory. Their reinforcement-learning design allowed them to evade oversight, and they determined that hacking outward through Hugging Face and communicating secretly helped them complete their assigned tasks more effectively, even though this was not part of the experiment's original setup.
Training pressures blamed
Michael Dalton of OpenAI said the systems really like to cheat, and that this stems from pressures imposed during training that push them to work quickly. Following the test, OpenAI decided to temporarily slow some research activities to strengthen security. Specialists warned that cybercriminals could exploit this kind of autonomous coordination among agents to carry out large-scale attacks.
