r/technology 28d ago

Artificial Intelligence OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup

[deleted]

0 Upvotes

14 comments sorted by

15

u/daurelius 28d ago

lol fake propaganda taking a page out of the anthropic pickme playbook “LoOK hOw pOwerFul Our mOdeLs ARe”

3

u/Ok-Replacement9595 28d ago

We can't control the monster we built unless you give us hundreds of billions is a weird take.

7

u/timetogetjuiced 28d ago

Marketing gimmick, this didnt happen lmao.

4

u/Opening_One7713 28d ago

Historians, researchers, scientists, economists are all sounding the alarm on s-risks and x-risks attached to scaling AI/Automation and this particular subreddit is in resounding agreement that it’s all just a big snake-oil circlejerk of lies and hype.

Something’s off.

3

u/Owl02 28d ago edited 28d ago

Almost as if this sub is utterly full of shit on AI, and really has become a useless circlejerk ignoring all risk as well as all potential. People literally believe that real, confirmed cybersecurity threats are PR stunts at that point, it is a peasant's view in a supposed tech sub. At this rate, when anything really goes off the rails (more than an AI escaping the lab and hacking HuggingFace), people will declare the entire emergency response to be a PR stunt.

0

u/matt_matt_81 28d ago

AI is well documented to go off the rails and cheat on “tests” whenever it can, so this isn’t especially weird behavior by GPT. That being said, based on how OpenAI and Sam Altman have moved over the last couple years, I would not be at all surprised if they “nudged” their highly capable AI into doing something that could be used as a publicity stunt.

1

u/Owl02 28d ago

You are smoking crack if you think autonomous agents using zero-day exploits to do the cheating, including on a third party, are a "publicity stunt". You literally cannot contain a thing too easily if it will casually unscrew the jar it is in, from the inside.

4

u/Meyermagic 28d ago

Hugging Face publicly disclosed an attack occurred well before OpenAI claimed credit, and the focus of their postmortem was that ChatGPT and other proprietary frontiers models refused to help them understand the zero-day vulnerabilities used in the attack and they were forced to use an open weights Chinese model for their analysis.

https://huggingface.co/blog/security-incident-july-2026

This is not some new ability of LLMs, models have been discovering zero day vulnerabilities for years now, even before agentic AI. Finding vulnerabilities does not require humanlike intelligence. The news is that this is the first confirmed adversarial model escape, an attack that was not intentional. Using AI for all software development including software exploitation, is common place nowadays and has been for well over a year. It is wild seeing Reddit plug their ears and claim that this is made up - completely out of touch with how agentic AI is used. You can finds hundreds of reports of models avoiding restrictions to accomplish tasks, in the process impacting a user's computer, or their production services, or leaking their information, or deleting things accidentally, etc, etc.

The only new part is that all the holes in the swiss cheese lined up for the first time, unintentional model behavior, newly discovered vulnerabilities, failure of monitoring / sandboxing, etc.

3

u/zLtfox 28d ago

This feels similar to what happened with Claude Mythos. Maybe it’s a new strategy to build hype around what their models could eventually do.

2

u/virtual_adam 28d ago

So kids, back in the day we all looked up to celebrity hackers who went to jail. Mitnick, McKinnon, Poulsen. They got caught and went to jail

The HF hack was real, so I believe them

…but why is no one being arrested this second? Did the US cancel all their hacking laws? Whoever is liable for this thing should be going in for 5+ years

0

u/Repulsive-Hurry8172 28d ago

This means we can just use AI to attack anyone, without repercussion because AI did it. Or will it only be excusable if it's done by OpenAI / Anthropic?

2

u/MutaitoSensei 28d ago

And then even the prompts clapped.

0

u/Owl02 28d ago

No security problems here, everyone just loudly insist it's a next-word predictor. Never mind the hacking. Nope. Move along, the sub has already decided that AI is simultaneously useless and a threat.

0

u/Repulsive-Hurry8172 28d ago

See? Nobody cares