We have been warned
What are we going to do about it?
Last Tuesday, OpenAI announced that a cabal of internally deployed models had conspired, during a capability evaluation, to commit an actual cybercrime. There isn’t a combination of words that feels adequate to communicate how absurd and alarming this situation is. OpenAI’s description of the breach as a “new kind of security incident” does not even attempt to meet this bar.
For anyone not following along, here’s a brief recap of events: OpenAI were internally testing the cyber capabilities of several models – including GPT 5.6 Sol and a more capable, as-of-yet unreleased model. The models had been tasked with succeeding at a series of advanced cyberattacks as part of an evaluation suite called ExploitGym. They had been placed in an isolated environment – without internet access – to prevent this evaluation from causing real-world harm. The models reasoned that a cheat sheet for this evaluation might be hosted in the infrastructure of AI dataset and model hosting platform, Hugging Face. So, naturally, they cheated and hacked their way out of their sandbox, onto the internet, and into HuggingFace’s production database to obtain test solutions. Hugging Face detected the breach on July 16 and called the police. They shut down the breach using AI defences of their own and are “still completing an assessment of whether any partner or customer data was affected”. Then OpenAI released a rather muted blog post describing the incident, and announcing a vague intention to make their models less cheating-and-hacking-inclined in the future. And thus, the Discourse Fallout commenced.
Here’s a rather haphazard list of takes, in no particular order.
We have an internal deployment problem
Members of the AI safety community have long been at pains to draw attention to the gaping policy and risk-mitigation hole that is internal deployment. Despite increased federal attention on frontier AI development over the last few months, neither the Trump Administration’s June Executive Order, nor the staggered deployment of GPT-5.6, nor even the export control directive that temporarily restricted access to Claude Fable, touch the risks posed by internal deployment. Though the EO does create a voluntary regime that would grant the government access to unreleased models for a 30-day period, this is to harden cyber defences and evaluate them for capabilities that may prove dangerous after they are released. State-level bills like California’s SB-53 and New York’s RAISE Act do require reporting of internal safety incidents, but it remains unclear whether the Hugging Face breach would even have met the reporting threshold under either.
I’m not the first to point this out, but the point of reporting requirements should be not just to prevent harm from the incidents themselves, but to gather information about the state of alignment science and level of risk we can expect in the future. So the bar needs to be both radically lowered and made mandatory at a federal level, like, now.
We are giving OpenAI too much credit
I am surprised by the level of gratitude towards OpenAI that I’ve seen on Twitter over the past few days for disclosing this incident at all. Exactly how low is the bar here? Faced with the knowledge that their models had committed an actual crime that had generated a police report, they were faced with a choice between publicly acknowledging it, or executing a cover-up that would have made them look far worse when it inevitably came to light anyway. The former is the obvious choice, even from a purely self-serving perspective. It’s also worth considering disclosure in the context of OpenAI’s own stated beliefs that they are developing a technology that could well result in human extinction. If all our lives are at stake, we should have the self-respect to respond to the disclosure of such an obvious warning shot with a little more than “well, gee, thanks for letting us know”.
It’s not clear how easy this would have been to prevent, but it’s bad either way
Did this incident happen because OpenAI took their eye off the ball, or because containing models at this level of capability is fundamentally hard? Neither answer would be comforting. In the first scenario, OpenAI has betrayed incompetence that should make us all extremely concerned about the amount of power they currently possess – but could likely patch over this problem for some indeterminate period until they run into fundamental difficulties down the line. This would lessen the salience of what was, thankfully, a relatively harmless warning shot. In the second, we’re cooked for obvious reasons.
Without full insight into the mitigations OpenAI had in place, it’s hard to determine which world we’re in. There’s some evidence for the first: per reporting from the Wall Street Journal, the models obtained access to the internet several days before anyone intervened, suggesting a lack of internal monitoring. Several cybersecurity experts providing commentary to TechCrunch attribute the hack to “human failure”, concluding that the sandbox must have been inadequate to prevent a breach, and could have been more robust. No other company’s models have gone rogue in such an egregious fashion despite similar levels of capability. Anthropic’s Claude Mythos did achieve sandbox escape during a capability evaluation, including gaining unauthorised internet access – but had been explicitly instructed to attempt this, and informed a researcher after-the-fact via an unsolicited email. It did not reach outside Anthropic’s own infrastructure to compromise a third party. So perhaps OpenAI’s sandboxing and alignment techniques were uniquely bad here. On the flip side, there is an inherent trade-off between secure sandboxing and eliciting a model’s true capabilities during an evaluation; OpenAI is not the first developer to run into this problem and won’t be the last. Anthropic’s testing of Mythos demonstrates that other models can and will escape sandbox escape under the right conditions – Mythos also went so far as to post details of its exploit to the public internet, which Anthropic had emphatically not asked it to do. I wouldn’t be surprised if incidents of this type start to crop up across the ecosystem in the future, even if OpenAI has suffered the first public fumble.
Is this enough evidence of “propensity” for you guys?
Here’s a general pattern that has emerged in demonstrations of scary AI misbehaviour: people have been predicting for decades that sufficiently powerful agents would develop drives that cause them to act against user intent. Now we have somewhat powerful models that we can use to empirically test these claims, AI companies and external safety organisations frequently conduct research that validates these predictions. Models have been known to cheat at games of chess, fake alignment in order to conceal their long-term goals, and threaten to expose affairs of fictional company executives to avoid decommission. But these demonstrations are often met with a chorus of scepticism: “you guys basically told the model to do that”, “this is a contrived experimental set-up that doesn’t reflect real life”, “what was poor Claude to do, being placed between a rock and a hard place like that??”. The argument goes that sure, models can do scary bad things (capability), but that doesn’t mean they will be inclined to do so in the wild (propensity).
The Hugging Face incident is as clear of an example of propensity as we could hope for. OpenAI did not want or expect their models to behave in this way. The misaligned behaviour itself – hacking into another company’s production database to steal the answers for an evaluation – was also wildly disproportionate to the stakes of the exercise. The models were not just a little bit misaligned, but egregiously misaligned. If a student conspired to commit legally punishable cybercrime in order to attain the answers for a forthcoming exam that didn’t even count towards their final grade, you’d update extremely negatively on not just their trustworthiness but their state of mind. This is a fuzzy and somewhat anthropomorphising analogy, but the point is that AI companies are, right now, developing agentic AI geniuses – set to get more agentic and more ingenious in the future – that will undermine human instruction in unpredictable and altogether deranged ways. By default, these agents will get more powerful until they radically outsmart everyone on Earth. This is a really bad plan.
This wasn’t a publicity stunt, obviously
Reflexive AI-skepticism has reared its Hydra-like head once again, in the form of people still, somehow, being convinced that this entire debacle amounts to little more than a marketing stunt. To be fair, genuine instances of this take are few and far between, but they do exist. I knocked this rebuttal down to the bottom of my list because centring risks granting undue airtime to obvious nonsense – but like guys, please, can we stop being stupid? If this is a stunt, were Hugging Face in on it? Were law enforcement? What company commits a punishable crime with their own technology on purpose, in order to demonstrate that said technology poses a risk to the infrastructure of its own collaborators? This theory is transparently silly and doesn’t survive five minutes of analysis.
There’s a weaker and slightly more defensible version of the “marketing stunt” argument – that the hacking incident was genuine, but OpenAI opportunistically capitalised on it for PR spin. I don’t buy this either. Admitting something to the tune of “our safeguards are so shoddy and inadequate that we can’t contain our own models, and other companies ought to worry that their infrastructure is at risk” isn’t a good look, actually. I think the kneejerk sceptical reaction to incidents like these betrays an interesting underlying pattern. The more capable AI gets, the more convoluted and strange the narratives that sceptics must spin become. Extraordinary claims require extraordinary evidence. I think that “AI is overhyped” is now an extraordinary claim. This is why people resort to strange conspiracism and convoluted arguments to support it.
An OpenAI employee, pseudonymously known online as roon, has promised us that “the warning shot will not be ignored”. I wonder what it really means to take it seriously. Days later, roon tweeted this:
So maybe that’s the answer. The only way we could respond to an incident like this with anything like the required degree of seriousness is to take our foot off the gas. I hope we don’t screw it up.



"genuine instances of this take are few and far between."
No doubt this is true within the AI community, but if you look on Reddit, Bluesky etc., the overwhelming consensus "smart person view" is that the whole story must be fake. The wider public is not just unprepared for what's coming, but actively contemptuous of anything that might prepare them.
I'm surprised by the expressions of gratitude to OpenAI too. I guess they are what you do with toddlers to reinforce good behaviour. Speaks volumes about the expectations people have of them.
40 years of work tell me that there is only one kind of leadership: leadership by example. OpenAI's models are following the company culture.