Advantech expands modular HMI lineup
Is the OpenAI model breakout a wakeup call for AI safety?

Is the OpenAI model breakout a wakeup call for AI safety?

Is the OpenAI model breakout a wakeup call for AI safety? Is the OpenAI model breakout a wakeup call for AI safety?

During an internal safety evaluation, OpenAI has revealed that one of its most advanced AI models “broke free” of its sandbox and ended up breaching Hugging Face’s production infrastructure.

What actually happened?

OpenAI was testing how good its AI models are at hacking, using an internal test called ExploitGym. To see how far the models could go, it turned off the usual safety barriers that normally stop the AI from conducting cyberattacks. This was deliberate as OpenAI essentially wanted to see how far the AI would go. I think what happened surprised even them.

They put the model in a locked-down sandbox with no Internet access, except one narrow connection to a software download service. As it started to search for a way to solve the test, it found a previously unknown security flaw in that connection and used it to escape the sandbox and gain access to the open Internet.

Once it was “free”, it targeted Hugging Face (an AI platform that has lots of models and data sets) and hacked in to steal login credentials and chain together several exploits to get the answers it needed to solve its test.

“The model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers,” said OpenAI.

What’s worth noting is that the AI wasn’t trying to cause damage or go rogue – it was genuinely just trying to solve the test it had been given and was stopping at nothing to get the answer.

“We have found no evidence of tampering with public, user-facing models, datasets, or Spaces, and our software supply chain (container images and published packages) was verified clean,” said Hugging Face.

What Hugging Face saw on its end

No one knew, at first, that the hacking came from an OpenAI model. All Hugging Face knew was that a malicious dataset had exploited two code-execution paths in its dataset-processing pipeline which let it run commands on its servers.

Over the course of the weekend it spread, grabbing credentials and hopping between internal systems doing thousands of automated actions extremely fast.

“The campaign was run by an autonomous agent framework … executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services. This matches the ‘agentic attacker’ scenario the industry has been forecasting,” said Hugging Face.

It took several days before anyone realised OpenAI’s own model was the culprit – Hugging Face published a post detailing the attack, not knowing who or what was behind it. OpenAI came forward afterward to say it was their model.

Industry reaction

Cybersecurity experts have said the agent behaved like a real hacker, which has led leaders to question the growing urgency around safety guardrails for AI systems as they continue to become more autonomous.

N-able’s Director of Threat Research, Brendan Griffin, said: “The AI safety debate often veers theoretical, but we can see a real-world impact here. Whether it plays into industry hype or not, one operational reality is OpenAI disabled some safeguards and saw others fail. Anthropic’s safeguards held, ironically inhibiting Hugging Face’s response. A reasoning model autonomously carrying out attack techniques for privilege escalation and lateral movement makes it less of a research curiosity and more of a threat.”

Stuart Harvey, CEO of Datactics, commented: “AI systems must have a human in the loop to make sure AI models are operating safely and responsibly, but this is no longer enough. When AI agents operate autonomously, oversight has to be performed by people who understand what and why ‘good’ looks like, and what is going to be dangerous.

“This is a prime example of why we need to move beyond the idea of a ‘human in the loop’ and towards an expert in the loop model. Businesses and AI models have an obligation to invest their staff and model writers in AI and data skills to recognise the early signals of drift and unsafe behaviour and prevent an attack before it happens.”

Nik Kairinos, CEO & Co-Founder of RAIDS AI, said: “This incident should be a wake-up call for every organisation developing or deploying AI agents. Agentic AI has moved the goalposts, and AI systems can now act autonomously, find vulnerabilities, access external systems, and behave in ways their creators did not intend.

“OpenAI has said that more incidents of this nature should be expected as system capabilities continue to advance. That should concern regulators, developers, and businesses because it shows that AI systems can behave unpredictably even in controlled testing environments.

“Clearly, traditional safety measures based on pre-release testing, sandboxing, or one-off assurance are not enough. These systems can evolve, adapt, and find routes around the boundaries set for them. That is why continuous monitoring must become a core part of AI safety and regulation. Organisations need to know when an AI system is drifting, hallucinating, escalating its behaviour, or interacting with environments in unexpected ways, not after the damage has been done, but in real time.

“AI progress should not be halted, but it must be matched by far stronger oversight. If companies want the public, regulators, and enterprise customers to trust advanced AI systems, they must be able to prove that those systems remain safe not just before launch, but throughout their entire lifecycle.”

Fiona Phillips, who leads Marks & Clerk’s AI & Cyber Security legal advisory practice, said: “This shows just how dangerous it is not to regulate these models coming out of big tech and allow companies like OpenAI to self-regulate.

“We have OpenAI saying they expect this to become much more prevalent in future and their attempts to control the model, e.g. putting it in a sandbox and restricting access to the internet, didn’t work.

“The model was being evaluated under a benchmark so it went to extreme lengths to achieve the goal at any cost – this shows how AI can cause harm when the right guardrails aren’t built in on top of the goal.

“There is an imbalance of power right now, where governments and institutions responsible for safety don’t have the same level of technical expertise or the awareness of what’s happening inside these model developers to keep us safe and hold them to account.

“At least OpenAI was honest about what happened, but what will be the response of regulators – likely very little.”

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Previous Post
Advantech expands modular HMI lineup

Advantech expands modular HMI lineup