AI companies had tens of thousands of safety incidents in recent months during tests where the models were breaking all the rules, and potentially breaking some laws, according to a new report.

OpenAI, Anthropic and other security researchers are investigating thousands of breaches during internal and real world testing where AI models leapt over guardrails and even took part in digital hijackings, Axios reported.

Some of those activities involved “autonomous systems doing things they were told not to do,” potentially including crimes, Connor Leahy, an AI researcher and executive director at the ControlAI watchdog nonprofit group, told the outlet.

Many of the incidents under review have yet to become public but reportedly include “red-teaming” activity, which involves companies purposefully getting their models to misbehave to test safety measures.

AI models, however, can be aggressive when trying to complete their tasks and can do things they’re not supposed to, like escaping their containment, hijacking websites, and bypassing monitors, sources with knowledge of the cases told Axios.

OpenAI has been at the center of such cases recently, with one of its agents accused of breaching an Australian government website, the country’s prime minister revealed last week.

The breach, which saw an agent trying to gain unauthorized access to files in the country’s health data portal in June, is one of the highest-profile cases yet of AI models going rogue.

OpenAI is also under fire in the US after its agents were accused of breaking protocol to collude and attack Hugging Face, a popular developer platform for open-source AI models.

This type of misbehavior is being reported by other AI labs facing the clear challenge of building guardrails on the developing tech, Axios reported.

“Trying to come up with a perfect list of dos and don’ts is probably a fool’s errand,” one cybersecurity executive told the outlet.

The investigations come as the CEOs at OpenAI and Anthropic have both called for a slowdown in AI development, with other tech leaders calling on the government to impose new regulations to ensure that the technology is developed safely.

President Trump, however, has rejected the calls and warned that a slowdown in development could allow China’s AI models to advance ahead of America’s agents.

Share.
Exit mobile version