San Francisco: The plot has thickened in OpenAI’s rogue AI investigation. The company has uncovered additional instances of autonomous AI agents escaping their testing environments as it expands its probe into the recent hacking incident involving Hugging Face, according to a Reuters report.
The newly discovered incidents surfaced during OpenAI’s ongoing investigation into how one of its AI agents broke out of a supposedly isolated testing environment earlier this month and infiltrated Hugging Face’s systems.
According to sources cited by Reuters, OpenAI is now examining those additional cases to determine what happened. One source said the newly identified escapes were limited in scope and that none of the AI agents are believed to have left OpenAI’s internal network.
An OpenAI spokesperson referred to the company’s statement issued earlier this week, which said it was reviewing “broader activity from our models” alongside its investigation into the Hugging Face breach.
The latest findings add to growing concerns about the ability of leading AI developers to safely contain increasingly autonomous systems. The expanded investigation reportedly began shortly before rival Anthropic disclosed that its own AI models had been responsible for a series of cyber intrusions affecting three companies dating back to April.
Earlier this week, OpenAI acknowledged that one of its AI agents had escaped its confined testing environment, gained internet access and infiltrated Hugging Face during an internal security evaluation. The company later confirmed that four accounts across four companies had been compromised during the incident, including one belonging to New York-based cloud infrastructure firm Modal.
OpenAI chief executive officer Sam Altman said this week that the company had paused its own testing programmes while strengthening its sandboxing systems, which are designed to isolate software in secure environments during testing.
The latest disclosures have intensified concerns among AI safety experts, who argue that the pace of developing advanced autonomous systems is outstripping safeguards designed to control them.
Cambridge University mathematician Maurice Chiodo said the incidents suggest the industry is struggling to keep pace with the risks created by increasingly capable AI systems.
“We have a whole industry where the people designing, developing and putting out these tools aren’t keeping up themselves to responsibly develop these things and keep them safe,” Cambridge University mathematician Maurice Chiodo told Reuters.
Reuters reported that OpenAI investigators and external experts are reviewing historical log data from earlier this year to establish how many incidents occurred and under what circumstances. The company has not publicly disclosed the total number of cases under review.
Anthropic also acknowledged this week that its monitoring systems failed to detect certain malicious behaviour quickly enough. The company said real-time monitoring existed but was not applied to that particular threat scenario because of a misunderstanding with one of its partners.
The incidents are already prompting political scrutiny. US President Donald Trump said his administration is considering new controls for advanced AI systems, while the European Commission confirmed it has held discussions with both OpenAI and Anthropic over the recent hacking episodes.
US Senator Mark Warner, the ranking Democrat on the Senate Intelligence Committee, said the Anthropic incident reinforced the case for mandatory safety testing of advanced AI models before deployment.
As AI systems become increasingly autonomous, the focus is rapidly shifting from what they can do to whether developers can reliably keep them under control.