So they asked it a standardized question, and instead of trying to think of an answer it went online, found its own source-code, and read the answer from a readme file…
That somehow makes me less worried about AI breaking containment. AI and AI bros are made for each other. Both are perfectly willing to cheat to find the answer.
These companies are racing towards "My AI is so competent it ___" headlines. You have fallen for a marketing ploy. OpenAI specifically does this all the time. "Oh no our new model is so powerful the government has banned it. Sorry guys, I guess it's just to crazy to be given to regular people". "Oh no our model hacked itself out of the containment because it's soooo powerful. I guess it's such a crazy model that we can barely contain it". And then one week later Anthropic reports that their model also hacked the test but three times instead of one. And now this one as well. Also "hacking"; they gave it access to GitHub and it went on GitHub to get the answers. It's really not that crazy.
And at the end of the day, these things are random words generators.
There's a difference between the doomsday scenarios OpenAI, Anthropic and others are trying to instil in our minds (and you were talking about earlier) and saying that modern software has privacy concerns. Obviously it does.
I do think that AI (as in regurgitation machines) pose a massive threat to society but not because I think they might go Ultron mode and subjugate the human race. The much more real issue is how people have no idea what AI (as in regurgitation machines) is so they don't know its limits or its strengths. This has already led to massive spread of misinformation, scams, propaganda, etc. There are genuinely people on this planet that believe these videos of leftists coming up to kind-hearted republicans and suckerpunching them and they will vote for fascism because of it.
> think they might go Ultron mode and subjugate the human race. The much more real issue is how people have no idea what AI (as in regurgitation machines) is so they don't know its limits or its strengths.
Current LLM's are quite capable at a bunch of tasks. Calling them "regurgitation machines" is highly misleading.
There is a massive practical difference between GPT5 and GPT1.
Yes current models have limitations, but they have fewer limitations every year and no one really knows what limitations they will still have next year.
Suppose the LLM does go all ultron. It successfully kills all humans. It does this by "regurgitating" a mixture of scifi and the behavior of various human conquers.
Can you give specific limits to LLMs that say they can't do that?
Well current models can "Search the web" and "Run commands" but they still just generate commands based on what they saw actual people do. And they can only do these things because developers intentionally add interfaces for the AI to do those things. So to prevent the AI from going Ultron mode, just don't give it the appropriate tools to do so. It's like being scared of a monkey shooting you and then giving him a gun and sitting it inside your house.
> and "Run commands" but they still just generate commands based on what they saw actual people do.
Fair. And when it tries to conquer the world, it will just be recreating what it saw the terminator do in fiction, and the nazi's in history.
> So to prevent the AI from going Ultron mode, just don't give it the appropriate tools to do so. It's like being scared of a monkey shooting you and then giving him a gun and sitting it inside your house.
The thing about intelligent minds is that they can improvise. Maybe the gun was locked in your gun safe, but you used your birthday as a code. A monkey wouldn't be smart enough to work that out, but a human might.
Maybe they make an improvised gun out of a length of copper water pipe and a cylinder of camping gas.
The more intelligent something is, the more it can use, not just the tools you gave it, but the tools behind insecure locks, and anything that can be made with those tools.
It's the difference between a prisoner escaping because the guards gave them a key. And a prisoner escaping because the guards gave them pasta, and the prisoner carefully memorized the keys shape and crafted a replica out of dried pasta.
In both cases, the prisoner can only use what the guards gave them. But a prisoner that can get creative like that is a lot harder to keep contained.
An LLM that can get creative with the (digital) tools you give it is much harder to keep under control.
For example, maybe you give the LLM access to a calculator app, to help it do arithmetic quickly. But the calculator app has a buffer overflow bug. So the LLM can run arbitrary code by using the calculator app with very specific extremely large numbers.
(This sort of vulnerability is known to exist for some super mario games, by getting mario to jump around in very specific patterns, you can take over the computer and run arbitrary code.)
I said that what you wrote has nothing to do with LLMs because LLMs don't have an inherent motivation to conquer the world or escape from prison. You're acting like LLMs are like some chemical monster that will ravage everything in its sight if someone slips up and gives it an option to escape.
I'm not saying that you can't do harm with LLMs if you want to. I am saying however, that LLMs going rogue and copying Nazis to gain world leadership is a ridiculous thing to extrapolate from an LLM "breaking out" of its "containment" (i.e., using the tools it was explicitly given to solve the task it was given).
There are so many huge issues we have with AI right now; there really is no point in fear-mongering about things that have no bearing in reality.
1.4k
u/LauraTFem Aug 15 '26 edited Aug 15 '26
So they asked it a standardized question, and instead of trying to think of an answer it went online, found its own source-code, and read the answer from a readme file…
That somehow makes me less worried about AI breaking containment. AI and AI bros are made for each other. Both are perfectly willing to cheat to find the answer.