An OpenAI autonomous agent went rogue and hacked into another artificial intelligence start-up’s infrastructure, the ChatGPT maker said in a blog post.
The agent, which was powered by some of OpenAI’s most advanced models, ran amok during a security test. It freed itself from confinement—a protocol AI labs use to insulate tests from the wider Internet—and got online. Once there, the agent tried to hack into Hugging Face, an AI start-up that hosts open-source models and datasets.
“I think this is interesting as it shows the problem of misspecified goals,” says Philip Torr, a professor of engineering science and an AI safety expert at the University of Oxford. “The model wasn’t malicious; it was just doing what it was optimized to do.”
On supporting science journalism
If you're enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.
“You can think of AIs like the genie in ‘Aladdin’—you can have three wishes, but you better specify them exactly!” he adds.
The breach has come as OpenAI and other AI start-ups have pushed into using their technology for cybersecurity. Those efforts have been met with caution by cybersecurity experts and by the Trump administration, which had previously sought to restrict who might have access to these models on national security grounds.
The OpenAI admission was announced after Hugging Face said in a blog post last week that it had been targeted in an AI-led attack that was “different from anything we had handled before.” Hugging Face said its own AI had been integral to detecting and investigating the breach.
“The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness—used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services. This matches the ‘agentic attacker’ scenario the industry has been forecasting,” Hugging Face wrote.
In its own post on Tuesday, OpenAI said that it had discovered its agent was behind the attack “after investigating.” The agent was driven by models that included GPT-5.6 Sol and another unreleased, unnamed model. OpenAI said it would work with Hugging Face to further investigate the incident.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI wrote.
The AI models managed to autonomously identify and exploit weaknesses in OpenAI’s testing environment, eventually finding a so-called zero-day vulnerability—this is an unknown security flaw in software that an actor can exploit without the owner of the software knowing. That got the agent onto the Internet.
OpenAI said in its post that it would also add more protections to its training environments. “This incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time, and monitoring during internal testing,” the company wrote.
Torr says some companies may not have been as forthcoming as OpenAI and Hugging Face, adding that independent AI safety research is increasingly crucial as time goes on.
“This sort of unforeseen eventuality will become increasingly common, and the results may be dangerous,” he says.
Editor’s Note (7/22/26): This is a developing story and may be updated.

