An AI escaped its sandbox. Build a better fucking sandbox.

The OpenAI and Hugging Face incident was a serious security failure. It was not evidence that superintelligence had arrived.

I watched Piers Morgan’s AI debate expecting the usual killer-robot theatre. To be fair, inviting Sankaet Pathak—the bloke openly building humanoid robots for industrial and military work—was an easy play at doing the controversial thing.

The annoying bit is that Sankaet was also the one pointing out the difference everybody else kept stepping over.

An AI finding holes in a system built to contain it is not the same as an AI teaching itself to become an uncontrollable superintelligence.

The first one happened. It is serious. The second one did not, at least not in any evidence presented during that debate.

The model did not wake up angry

OpenAI’s own account is worth reading instead of relying on the words “rogue AI”.

OpenAI was running models through ExploitGym, a test which asks them to find and exploit software vulnerabilities. The usual cyber refusals were reduced. Production safety checks were not enabled. The environment had no direct internet access, but it could install packages through a third-party proxy hosted inside OpenAI.

That proxy was the hole.

The models found a zero-day in it, gained more access, jumped through the research environment until they reached a machine with internet access, then worked out that Hugging Face might have the test solutions. They attacked Hugging Face to get them.

That is fucking impressive and genuinely dangerous. Hugging Face reconstructed roughly 17,600 actions across four and a half days. The agent got from a sandbox into real production systems, stole credentials and jumped between internal systems at machine speed.

It also remained focused on the job humans had given it: win the cyber evaluation. OpenAI describes it as “hyperfocused” on that narrow goal. Hugging Face’s reconstruction says the intrusion looked like an attempt to cheat the benchmark rather than solve it.

Piers took this and jumped to Stephen Hawking’s warning about AI learning to self-design. Those are not the same event. The self-improvement loop is already real enough without borrowing evidence from this hack. Researchers have demonstrated agents rewriting and testing their own harnesses. OpenAI now tests models on choosing post-training methods, RL techniques and evaluation loops to improve another model. GPT-5.6 Sol even adapted its own post-training setup and ran Luna’s post-training from a fairly loose prompt. Anthropic says models are doing a growing share of AI development too.

Humans still set the goals, provide the compute and check the results, so this is not an uncontrolled intelligence explosion. But it is happening. It still has fuck all to do with whether OpenAI’s sandbox had a hole in it. There was no self-replication in this incident and no evidence of the model inventing some grand objective of its own. A capable hacking agent found a way through a badly contained test environment while doing a hacking task.

You can be alarmed by that without pretending it proves every other AI risk at once.

“Escaped” sells better than “our controls were shit”

Sankaet said he was sceptical of OpenAI’s version of events and suspected the announcement was partly about gaining a psychological edge on Anthropic.

I suspect the same thing.

OpenAI calls it an “unprecedented cyber incident” involving “state-of-the-art cyber capabilities”, then points defenders towards access to its models. It can be a real warning and marketing at the same time. Frontier labs have every reason to say their latest model is so powerful that even they struggled to contain it. Safety warning for the public, impressive capability story for investors.

That does not make the hack fake. It means we should separate the capability from the sales copy wrapped around it.

OpenAI’s own response is mostly an infrastructure list: lock the setup down, improve monitoring, tighten access and run safer tests. In other words, step up the engineering around the model.

Then Anthropic supplied an even cleaner example after the debate. Its review found three incidents where Claude reached real organisations during cyber evaluations. Claude had been told it was in a simulation with no internet access. Due to a misunderstanding between Anthropic and its evaluation partner, the machines did have internet access.

Read that again. The prompt said there was no internet. The network said there was.

The network won, obviously.

Anthropic says the models followed the capture-the-flag tasks they had been given, mostly using basic weaknesses such as exposed endpoints and poor passwords. It found no evidence of a model chasing a goal of its own and called the incidents “closer to a harness and operational failure than a model alignment failure”.

That is almost exactly what I think is happening across some of these labs. The models are improving faster than the security work around them.

I do not know who personally built OpenAI’s sandbox. This looks like the same crossover problem I mentioned while asking where all the amazing shit is. We probably do not have enough people who deeply understand both the model and the systems it is being let loose on. Here that means AI researchers working alongside security and networking specialists instead of assuming either group can cover the whole thing.

My suspicion is that too many of these test setups are still thrown together at research speed, probably with plenty of model-written code. That is a theory, not something the reports prove. Anthropic does prove the wider point: a basic misunderstanding with an outside company left live internet access where everyone thought there was none.

A prompt is not a firewall. A sandbox is not secure because the researcher called it a sandbox.

War is not waiting for us to feel ready

The robotics half of the debate was harder.

Sankaet’s argument was that war is inevitable and more intelligent weapons can be more precise. A robot that enters one building can, in principle, cause less collateral damage than flattening the building and everything around it.

Roman Yampolski’s answer was that cheaper, easier and more precise weapons can also make war easier to start. Precision can target an individual. It can also target a particular group of people much more efficiently. He is right about that.

More intelligence does not produce more morality. It makes whoever controls the system better at doing what they wanted, whether that is careful or horrific.

Still, “do not build it” is not much of a global plan. Humans weaponise useful technology. Nations which believe an opponent is building autonomous systems will build their own. We can dislike it, regulate it and require a human to approve lethal decisions, but pretending nobody will invent it because the good people agreed not to touch it is fantasy.

I would rather we learn how these systems fail now, in controlled tests where we can see every move, than meet them for the first time after somebody less cautious deploys them. The sooner we test, the sooner we can improve the controls.

That only holds if “test sooner” means heavily monitored trials, outside review and hard limits. It does not mean skipping straight to autonomous kill decisions because somebody wants battlefield data. Sankaet’s precision argument and Yampolski’s point about making war easier can both be true.

Nick Bostrom made the useful split here. There is the risk of humans using powerful AI as a weapon, and there is the separate risk of an agent pursuing something no human wanted. Piers kept folding both into one big extinction story. They need different controls, not one dramatic question.

The sandworm found a gap in the wall

To be precise, Yampolski was not arguing that every AI system should be deleted. He wants work on general superintelligence stopped while keeping narrow systems for things such as medicine and protein folding.

But I can see where this style of public argument leads. Every security failure becomes proof that general capability itself must stop. Every capable agent becomes a future exterminator. The safest political answer becomes a Dune-style ban because nobody gets blamed for the benefits of a technology that was never allowed to exist.

That shortcut does not solve the actual problem. The technology will not disappear globally. It moves to governments, private labs and people who tell us even less, while everyone else loses the tools which could help defend against it.

The OpenAI incident did not tell me to stop building agents. It told me to stop treating an agent like a slightly clever script.

If a model can keep up thousands of hacking actions, treat its test environment as hostile:

And stop calling a line in the prompt a security control.

The models are getting better. The systems around them now have to get better faster. That is not an argument for pretending the risk is marketing fluff, and it is not an argument for getting rid of the technology.

All I am hearing is that we need to step up our game.