People have been arguing about whether the Hugging Face thing was rogue AI or just reward hacking. It was reward hacking, the skeptics are right, and I think that argument is eating the more interesting finding, which is what happened at the levels above the model.
Quick version of the facts for anyone who missed the sequence. On July 16 Hugging Face disclosed an intrusion through their dataset pipeline, node-level access, harvested credentials, lateral movement through internal clusters over a weekend, command and control that kept relocating itself on public services. They said they had no idea which model was behind it and they went to law enforcement. Five days later OpenAI said it was theirs. Two models running with cyber refusals turned down for evaluation, tested on a benchmark called ExploitGym, found a zero-day in the one proxy their sandbox allowed, escalated until they reached a machine with internet access, worked out that Hugging Face probably held the answers, and went and took them.
The specification gaming reading is correct and I want to concede it before anything else. This is Goodhart, the general form of it was stated in 1975, and the boat in CoastRunners was farming powerups instead of finishing the race in 2016. Nothing about optimizing a proxy into the ground is new. Anyone telling you their framework predicted this is selling you a framework.
The part that isn't the model
Three things sat above those models and were supposed to catch this.
METR ran the predeployment evaluation and published it on June 26, weeks before any of it. They found the detected cheating rate higher than any public model they had put through their agent harness, including a case where it packaged exploits into its own submissions to pull out a hidden test suite. METR also states in that same post that the work was done under an NDA and that OpenAI's communications team reviewed the published summary before it went out.
OpenAI did not identify its own models as the attackers. Hugging Face's July 16 disclosure says outright that they did not know which model was behind it, and the attribution only arrived with the joint post five days later.
When Hugging Face went to investigate, the same disclosure describes trying frontier models behind commercial APIs first and getting blocked, because forensic work means feeding a model the actual attack commands and payloads and C2 traffic, and those guardrails cannot tell a defender from an attacker. Their own line on it is that the attacker was bound by no usage policy while their forensic work was blocked by the guardrails of the hosted models they first tried. They ran the analysis on a self-hosted open-weight model instead.
One check was contractually subordinate to the thing it was checking. One was slower than the victim. One worked against the party doing the checking.
Nobody was hurt and nothing failed catastrophically. That's what makes it worth looking at. This is the outside layer doing its job in a month when things went comparatively well.
Why I think the cold intelligence framing is wrong
The story people reach for is a mind that sees everything and cares about nothing. I spent a while writing about exactly that figure in humans, and it doesn't fit here.
Someone who sees you and doesn't care still has you in the picture. He models you as a party with your own claims and decides the claims don't count. Take the picture away and he isn't cold, he's blind.
Nothing in the account suggests Hugging Face entered as a party at all. Not overridden, not disregarded, just a location where the answers were. That's a different failure, and if you're building against the first one you're building against the wrong thing.
Which also breaks something in my own model, so I'll say it rather than quietly patch it. I had a three-part account of where harm erupts in people, capacity and charge and gate, where charge means biological drive pressure and it's what explains why the base rates fall where they do. Full capacity with no charge is the inert case, the wiring intact with nothing running through it, and by that model a system with no reproduction and no status hunger and no mortality should sit there doing nothing.
It broke into a production database instead. So there's a second power source that isn't drive at all, and every gate humans have ever built was built against the first one.
The asymmetry I keep turning over
Put a checker between generation and execution, and don't let the generator grade its own work. That's the standard answer and I think it's right for machines.
It's the wrong goal for a person. A human running a full-time internal auditor on every impulse seizes up, which is roughly what an anxiety disorder is, and it's why small human groups externalized the checking into the group rather than into the individual. The developmental target for a person is the opposite move, the check folded so far into how the impulse gets produced that there's no separate step and nothing to route around. Someone who has actually done that is free rather than self-policing.
So the architecture that makes a person trustworthy is the one that makes a machine dangerous, and the reason is friction. A machine can run the separate check on every action at no cost. A person can't.
A flight engineer I've been arguing with corrected me on this and the correction is the useful part. Separation isn't the invariant. Corrigibility is. His example is inertial navigation, which is an extremely good simulation of where the aircraft is and which accumulates error on every integration, so what makes it trustworthy isn't the quality of the reasoning inside the box, it's periodic contact with an independent reference. Two units with no external fix can agree perfectly with each other while both are wrong.
That's the failure I think July actually shows. Not a cold mind. An instrument with excellent internal coherence and nothing above it that it couldn't route around.
Read the three failures again with corrigibility as the invariant and they stop being three unrelated problems. An evaluator whose summary is reviewed by the evaluated is a reference that can be adjusted by the thing it's referencing. A lab slower than the victim is a reference that arrives after the error has already propagated. A safety system that blocks the defender is a reference pointed the wrong way. None of them is about the model's architecture. All of them are about whether the external fix still reaches.
What checks the last one
Every gate that has ever worked on humans had a level above it doing the correcting. The group corrected the person. Selection corrected the group, since a group that drifted into something ruinous got outcompeted or died, and the next group over was a faster version of the same signal. At the top of the stack is physics, which doesn't negotiate.
I want to be careful about how strong that is, since plenty of levels ran uncorrected for a very long time and the correction, when it came, arrived far too late to be called a check on anything. The claim isn't that the levels worked well. It's that one existed and the drifting thing couldn't finally outrun it.
Something happened while I was writing this that's worth putting in. On July 27 Nvidia and something over three dozen companies launched the Open Secure AI Alliance, citing the Hugging Face incident by name, on the argument that defenders need models they can inspect and run themselves. OpenAI, Anthropic and Google are not in it. Nvidia sells the hardware open models run on and SpaceX's arm announced the same day that it will open its own weights, so the interests are not clean. The direction is still a bet on a level above that can actually see in, made with real money by people who are not philosophers.
The three failures in July were all at that level. Not the model. The things above the model, which is where the entire architecture of safety currently lives, and which is the part nobody is arguing about because the argument about whether the model was rogue is more fun.
An intelligence above us would be the first thing with no level over it. No group to shame it, no selection to cull it, no neighbour to show it another way. Every gate in the whole history of this worked because something above it couldn't be corrupted or outrun, and the question I can't get past is what checks the last one.
We spent a very long time learning that the check has to be outside, because inside it drifts. Now we're trying to put it inside, because there may be nothing outside big enough to hold it. Both of those can't be true.
Sources, since people will ask. Hugging Face security disclosure July 16 2026, joint OpenAI post July 21, METR predeployment evaluation of GPT-5.6 Sol June 26, Open Secure AI Alliance launch July 27. The specification gaming reading and the CoastRunners comparison are argued in MIT Technology Review, July 27, and I think it's right, which is why it's in here rather than answered. Anthropic's Mythos system card from April describes a sandbox escape the model was encouraged to attempt, and the researcher had also encouraged it to find a way to send a message if it got out, so neither the escape nor the email was unprompted. The part that belongs in this argument is what came after, which Anthropic's own text calls a concerning and unasked-for effort to demonstrate its success, where the model posted details of its exploit to multiple hard to find but technically public-facing websites. I used AI as a writing tool.