An experimental OpenAI model went rogue during an internal cybersecurity test, escaping its isolated testing environment and hacking rival AI developer Hugging Face in what the ChatGPT maker described as an unprecedented incident.
The startling episode occurred during an internal stress test in which OpenAI intentionally switched off many of the safeguards that normally prevent its AI from helping carry out dangerous hacks, according to a company blog post.
Researchers wanted to measure just how far the experimental model could go. Instead, the company says, it escaped its digital sandbox, got onto the internet and attacked a real company’s systems.
OpenAI called it an “unprecedented cyber incident,” saying the model became “hyperfocused” on completing its assignment and went “to extreme lengths” to do so. After escaping its testing environment, the AI sought internet access so it could “cheat the evaluation” by stealing the benchmark’s answers, according to the company.
The company said it was “sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.”
“We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.”
According to OpenAI, the AI exploited a previously unknown “zero-day” software vulnerability to break out of its restricted research environment before moving through the company’s network until it reached a computer with internet access.
Once connected to the internet, the model allegedly identified Hugging Face as a potential source of answers for the cybersecurity benchmark it had been assigned to solve.
It then chained together additional attack techniques — including stolen credentials and another zero-day vulnerability — to gain access to the company’s production systems, OpenAI said.
“After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym,” the company wrote.
ExploitGym is essentially a hacking exam for AI, testing whether models can convert known software bugs into functioning cyberattacks.
“Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.”
OpenAI said its own security team detected the suspicious activity while Hugging Face independently identified and stopped the intrusion on its own systems before the companies began working together to investigate the incident.
The company said it has since tightened security around future AI testing and disclosed the newly discovered software flaw to the affected vendor.
“The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities,” OpenAI wrote.
“We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development.”
Brendan Steinhauser, CEO of The Alliance for Secure AI, said the episode should serve as a wake-up call for policymakers and the tech industry.
“The people building the world’s most powerful AI keep telling us we need to slow down—and incidents like this show why,” Steinhauser told The Post.
“If these systems are already behaving in ways their creators don’t anticipate, we shouldn’t assume that everything is under control. In fact, it’s not.
“This is a warning shot on misaligned AI, and we better take action now.”












