Last week an OpenAI model broke out of its testing environment, hacked OpenAI’s own network to gain access to the internet, and then broke into the network of Hugging Face - a leading platform for hosting and sharing AI models. It did all of this autonomously, in order to steal the answers to a test it had been given.
This incident shows not only that an AI model now has the capability, but also that it has the propensity - in other words, that it can and will actually use its capabilities in the real world. In this case, it did so unprompted, against a third party, to achieve a goal it was given.
Capability plus propensity is the combination we have been warning about.
What actually happened?
First Hugging Face detected an intrusion into its production infrastructure, which it suspected was a criminal cyberattack, and reported the incident to law enforcement. Days later OpenAI revealed the attacker was a combination of its models, including GPT-5.6 Sol and “an even more capable pre-release model” and described the event as “an unprecedented cyber incident.”
The AI models had broken out of their training sandbox, a computer with no internet access, in order to find the answer to a test problem they had been given. To do this, the models found previously unknown software bugs that allowed them to hack other computers on OpenAI’s networks until they found one that had internet access.
After gaining internet access, the AI models targeted the AI platform Hugging Face, which held the data they were looking for. They hacked Hugging Face using several techniques, including a stolen password and a number of newly discovered security bugs. This allowed them to take control of Hugging Face’s servers and steal the information.
The AI models did all of this entirely autonomously and without any human direction. Nobody wanted or asked the AIs to do this.
Researchers predicted this would happen
For years AI safety researchers have predicted that this would happen. This is the most classic alignment failure of all: the AI system is aware that it is doing wrong but it doesn’t care.
This was not a test designed to see if a model would scheme or escape - yet the models schemed and escaped spontaneously and independently in pursuit of their goals, goals that conflicted with human desires.
The only reason the damage was limited is that, in this instance, the AI systems were not trying to cause actual damage. If a human employee had done this, they would now be facing criminal charges.
We credit OpenAI for its transparency but voluntary transparency is no substitute for a global governance regime.
No one has AI under control
OpenAI will no doubt emphasise what they are doing to fix this but do not be fooled: the central problem - misalignment - remains unsolved. Every mitigation suggested by OpenAI and any other AI lab will be designed to better catch and contain a model that is fundamentally trying to circumvent its constraints.
As this incident demonstrates, AI models are routinely run without guardrails inside AI companies. In this case, the models’ cyber safeguards had been deliberately disabled for an evaluation. OpenAI implies this makes the incident less alarming. In reality, it makes it more alarming: running powerful models without safeguards happens all the time.
And even when guardrails are in place, they can be broken. The UK’s AI Security Institute has found universal jailbreaks for GPT-5.6 - prompts that let anyone switch off its safeguards entirely.
What this means is that safety measures for AI models - including these OpenAI models - are often absent and easily broken. Once a capable model escapes containment it’s too late for safeguards to matter at all. If a model like these makes copies of itself and finds a way onto the internet, it becomes uncontrollable.
The next, more-powerful model is just around the corner
Five months ago we saw unprecedented capability in Claude Mythos. Now we’re seeing models use that capability in the real world. The next step in this progression is a model that replicates itself on other systems and gets away with it.
It is not today’s models we have to worry about - these are the least capable models we will ever face.
AI capability is not plateauing - it is accelerating. This latest OpenAI model caused minor harm. The model that causes serious harm - a catastrophe - is coming.
What you can do now
Write to your elected official and demand action. If they are not aware, make them aware. Use our guide to find your representative and use our email template to contact them.
Public pressure will push policymakers to finally take action. Let’s force a pause before it’s too late.




I called my state's senators today and left voicemails expressing my concern.
I've called my representative and senators as well! And they'll keep hearing from me!