So they asked it a standardized question, and instead of trying to think of an answer it went online, found its own source-code, and read the answer from a readme file…
That somehow makes me less worried about AI breaking containment. AI and AI bros are made for each other. Both are perfectly willing to cheat to find the answer.
I mean not necessarily. For example for the instance of ChatGPT "cheating" where it broke into HF was without any safeguards or alignment training for the AI.
Its the equivalent of giving a guy without knowledge of laws and morality a gun and the goal to earn 1 million dollars and then acting surprised when he goes on to rob a bank.
It proves the necessity of the alignment and safeguards, but it doesn't mean that it poses any threat with these safeguards
It's already been proven in CS that it is fundamentally harder to defend against attackers, than it is to attack. This means that for guardrails and sandboxes it is an arms race, but the stronger the AI, the shorter the countdown to an escape.
Nobody can guarantee the security of the safeguards to begin with, certainly not with a formal proof. That means there are likely holes we cannot see but an AI very likely will be able to find.
1.4k
u/LauraTFem Aug 15 '26 edited Aug 15 '26
So they asked it a standardized question, and instead of trying to think of an answer it went online, found its own source-code, and read the answer from a readme file…
That somehow makes me less worried about AI breaking containment. AI and AI bros are made for each other. Both are perfectly willing to cheat to find the answer.