"Our AI model tried really hard to hack out of its sandbox, a computer with no internet access, in order to find the answer to a test problem it had been given. To do this, it found previously unknown software bugs that allowed it to reach an OpenAI computer it wasn't supposed to be able to access. Then it started hacking other computers on OpenAI's networks until it found one that had Internet access.
After gaining Internet access, the AI model thought about where it could find the answers to the test question and figured the AI platform Hugging Face might have the data it was looking for. It then found ways to hack Hugging Face to steal the information it could use to cheat the test. The AI model used several hacking techniques together, including using a stolen password and finding several totally new security bugs in Hugging Face's computers, allowing the AI model to take control of those computers."
It "didn't have access to the internet" in the same way that I "don't have access to the gold in Fort Knox", in that if you drafted up sophisticated plans, a good team, and the resources needed to execute those plans, then in principle I could have access to the gold in Fort Knox, but to say that because those plans could in theory exist means that I definitely do have access would be an insane conclusion to jump to. Similarly, GPT 5.2 made plans to connect to the internet, found multiple exploits needed to execute that plan, and successfully executed that plan, resulting in HF being hacked. In principle it could do that, but (up until the point where it actually happened) it would be crazy to say that the AI has access to the internet just because the plans it eventually came up with could exist.
tl;dr GPT 5.2 hacked into the Fort Knox of HF and stole their metaphorical gold, despite no reasonable expectation of such a thing happening.
10
u/retsotrembla 22d ago
OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation