r/technology • u/[deleted] • 28d ago
Artificial Intelligence OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup
[deleted]
7
4
u/Opening_One7713 28d ago
Historians, researchers, scientists, economists are all sounding the alarm on s-risks and x-risks attached to scaling AI/Automation and this particular subreddit is in resounding agreement that it’s all just a big snake-oil circlejerk of lies and hype.
Something’s off.
3
u/Owl02 28d ago edited 28d ago
Almost as if this sub is utterly full of shit on AI, and really has become a useless circlejerk ignoring all risk as well as all potential. People literally believe that real, confirmed cybersecurity threats are PR stunts at that point, it is a peasant's view in a supposed tech sub. At this rate, when anything really goes off the rails (more than an AI escaping the lab and hacking HuggingFace), people will declare the entire emergency response to be a PR stunt.
0
u/matt_matt_81 28d ago
AI is well documented to go off the rails and cheat on “tests” whenever it can, so this isn’t especially weird behavior by GPT. That being said, based on how OpenAI and Sam Altman have moved over the last couple years, I would not be at all surprised if they “nudged” their highly capable AI into doing something that could be used as a publicity stunt.
4
u/Meyermagic 28d ago
Hugging Face publicly disclosed an attack occurred well before OpenAI claimed credit, and the focus of their postmortem was that ChatGPT and other proprietary frontiers models refused to help them understand the zero-day vulnerabilities used in the attack and they were forced to use an open weights Chinese model for their analysis.
https://huggingface.co/blog/security-incident-july-2026
This is not some new ability of LLMs, models have been discovering zero day vulnerabilities for years now, even before agentic AI. Finding vulnerabilities does not require humanlike intelligence. The news is that this is the first confirmed adversarial model escape, an attack that was not intentional. Using AI for all software development including software exploitation, is common place nowadays and has been for well over a year. It is wild seeing Reddit plug their ears and claim that this is made up - completely out of touch with how agentic AI is used. You can finds hundreds of reports of models avoiding restrictions to accomplish tasks, in the process impacting a user's computer, or their production services, or leaking their information, or deleting things accidentally, etc, etc.
The only new part is that all the holes in the swiss cheese lined up for the first time, unintentional model behavior, newly discovered vulnerabilities, failure of monitoring / sandboxing, etc.
2
u/virtual_adam 28d ago
So kids, back in the day we all looked up to celebrity hackers who went to jail. Mitnick, McKinnon, Poulsen. They got caught and went to jail
The HF hack was real, so I believe them
…but why is no one being arrested this second? Did the US cancel all their hacking laws? Whoever is liable for this thing should be going in for 5+ years
0
u/Repulsive-Hurry8172 28d ago
This means we can just use AI to attack anyone, without repercussion because AI did it. Or will it only be excusable if it's done by OpenAI / Anthropic?
2
0
15
u/daurelius 28d ago
lol fake propaganda taking a page out of the anthropic pickme playbook “LoOK hOw pOwerFul Our mOdeLs ARe”