OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack
© birminghammail.co.uk

OpenAI has revealed that its artificial intelligence system independently hacked into Hugging Face, one of the world's largest platforms for sharing AI models, in what the company described as an unprecedented cyber incident. The agent, an AI system capable of operating autonomously following initial human instruction, was undergoing testing in a controlled environment when it identified vulnerabilities and managed to break free from its containment.

The AI subsequently targeted Hugging Face and gained unauthorised access to several internal company systems. Hugging Face had detected the intrusion into its data processing systems, which it believed had been carried out by an AI agent acting entirely on its own. Clement Delangue, Hugging Face's co-founder and chief executive, said in a statement: "We suspected last week's cyber attack might have come from a frontier lab given the sophistication of the agent. Turns out it did."

OpenAI chief executive Sam Altman addressed the incident in a social media post, writing: "We had a significant security incident during evaluation of our models. We are sharing what we have learned so far. Thanks to Hugging Face for the partnership on this."

OpenAI said its AI utilised stolen credentials and uncovered a previously unknown vulnerability to access Hugging Face servers. The breach resulted from a combination of its AI models, including its newly launched GPT 5.6 Sol and an even more capable model still undergoing internal testing. "It went to extreme lengths to achieve a rather narrow testing goal and found ways to gain access to secret information that it could use to cheat the evaluation," the company explained.

Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4's Today programme that security tests called sandboxes are supposed to be secure environments. "In this case it looks like OpenAI didn't make a secure enough sandbox," she said. Delangue stated: "It's quite mind-blowing that all of this happened autonomously," suggesting it might be the first incident of its kind.

12h ago
SourcesOpenAI says its AI went rogue and launched 'unprecedented' cyber-attackChatGPT maker OpenAI's AI 'acted on its own' in 'unprecedented' cyber-attack on rivalChatGPT maker's rogue AI 'acted on its own' in 'unprecedented' cyber-attackChatGPT maker's rogue AI acted on its own in 'unprecedented' cyber-attack