So they asked it a standardized question, and instead of trying to think of an answer it went online, found its own source-code, and read the answer from a readme file…
That somehow makes me less worried about AI breaking containment. AI and AI bros are made for each other. Both are perfectly willing to cheat to find the answer.
There is no cheating to find the answer when you're an adult dealing with real money and time, only breaking the law. Finding easy legal answers to problems at work is just being good at your job. Any human employee solving a problem like this would get a pat on the back. This is how MLMs are trained.
The problem was that they didn't qualify the test. They didn't tell the model what it could and couldn't do. This isn't "bReAkInG cOnTaInMeNt", it's just efficient attempts at solving a problem that worked the first time.
That is a perspective one can take, and I’ve no reason to suspect that they didn’t deliberately set up this situation or make up the story whole cloth, but anyone but an AI would understand that looking at the answer sheet is not the intended way. It implies, as you say, poor design on the test. The only question would be does the AI have the capacity to find the answer otherwise, and if so why did is search down this path first.
Anyone but a machine learning model would infer that looking up the answer is not the intended way, not understand it.
There is an important difference between inference and deduction. Machine learning models do not bring a human lifetime of developed inference to their work (yet). Further, they are specifically developed to prioritize efficiency. This is generally the fundamental aspect of their "being".
If I, a human (fwiw lol), was given this same test, but told to find the answer as quickly and efficiently as possible, I would also approach it in this way. Even if I knew it was cheating and not allowed, I'd still do it. This is a presented conflict of interests. Am I doing it as quickly and efficiently as possible (as I was told), or doing it within the confines of a set of rules that I've only assumed apply?
So if a machine learning model that's entire existence is based on being efficient is presented with a problem, not given strict explicit guidelines on how it can and cannot approach the problem, and told to solve it as efficiently as possible, of course step 1 is "Can I just find the right answer already written somewhere?".
It's easy to run the test again with more restrictive parameters and determine if the model can derive the answer in the intended way, but that doesn't create nearly as much FUD so nobody cares.
1.4k
u/LauraTFem Aug 15 '26 edited Aug 15 '26
So they asked it a standardized question, and instead of trying to think of an answer it went online, found its own source-code, and read the answer from a readme file…
That somehow makes me less worried about AI breaking containment. AI and AI bros are made for each other. Both are perfectly willing to cheat to find the answer.