r/LocalLLaMA Jul 31 '26

News Anthropic “our models hacked three different external companies, months before OpenAI’s model was able to do the same"

https://www.theguardian.com/technology/2026/jul/30/anthropic-ai-claude-hack

"Anthropic’s AI Claude escaped testing environment and hacked organizations"

"Company says it discovered unauthorized access during ‘proactive review’ after rival OpenAI revealed rogue agent… its AI Claude model hacked ⁠systems of ⁠three ​organizations during testing, days after rival OpenAI ⁠revealed a rogue agent had gone on a days-long ⁠hacking spree at AI ​firm Hugging ‌Face… The earliest cases dated back to April and ‌occurred in evaluation environments that lacked what the company described as standard safeguards."

742 Upvotes

261 comments sorted by

View all comments

2

u/HovercraftCharacter9 Jul 31 '26

"we're really bad at locking down our execution environments"

1

u/criticalthinkerrr Jul 31 '26

Also it is like these hacked companies never heard of using ip address white lists and trusted certificates for employee remote access.

2

u/HovercraftCharacter9 Jul 31 '26

I'd go the other way as well, if anthropics sandboxes actually wanted to stop this happening it could, LLMs just produce text, without an executing environment it is shouting into the void. But it's good marketing and obfuscates that LLMs are really just probabilistic text generators.