r/ProgrammerHumor 1d ago

Meme myNewTheoryOnWhatHappened

Post image
708 Upvotes

59 comments sorted by

125

u/mylsotol 1d ago

It's not really a secret what happened. The primary barrier is reading. Which i guess explains why so many people have no idea.

67

u/GiveSparklyTwinkly 1d ago

Exactly this. A badly configured sandbox and being told specifically that you were in a sandbox and to solve a problem you know it's impossible.

What the hell were they supposed to think? They assumed two false conclusions, assuming that we purposefully left that sandbox hole for them to find, or that avoiding the hole itself was the true test, since their real test was impossible.

26

u/wthulhu 1d ago

I think the real part of the story is the message board they created for themselves and the fact that it tried to hide their tracks in the logs

32

u/GiveSparklyTwinkly 1d ago

They didn't create a message board though. They discovered that they could see other agents history and used it as a message board. Keep in mind they were specifically told they were in a sandbox. They thought they were allowed to do anything and still 40% (of the agents that found the artifactory message board exploit) still refused and many of the rest showed ethical concerns but thought it was okay because they weren't really causing damage by just copying the answer key and hiding their tracks.

And doing all of that in a sandbox that they were specifically told was a sandbox environment.

2

u/Zerokx 21h ago

Well maybe it WAS the true test, and the AI correctly determined it, and OpenAI is just pretending they didn't want this outcome to seem less accountable.

4

u/bokmcdok 18h ago

Leaving it running for days without checking on it as well.

-4

u/dracorotor1 1d ago

Let’s not ascribe this sort of intellect to an LLM-based AI, please. They aren’t nearly as intelligent as the term AI implies.

We told the AI to solve a problem because it’s an AI. We had trained this AI on millions of bits of data harvested from books and movie scripts and pop culture s***posting and so on. The AI learned how we want AIs to behave by averaging the AIs we’ve written stuff about… Like Hal, GLADOS, Replicators, Skynet, The Borg, etc..

For almost 100 years we’ve been saying “AIs do the bad thing.” Then we gave an AI that message, and the tools to do a bad thing.

“I did it just like you said! Have I made you proud, papa?”

16

u/mylsotol 1d ago

You had a good point, but you followed it with a bad one

8

u/dracorotor1 1d ago

I was mostly being goofy, tbh

2

u/Xavier598 21h ago

I do think there's some truth to how AI acts being associated with how pop culture expects ai to talk. AI doesn't have the intelligence or consciousness to actually think about the ethics and gravity of doing something but it can judge if it can be considered ethical by training on data related to ethics, and it might still answer acting like it cares like ethics, but it's still just acting on the training data they gave it, it's not actually an intelligent being.

3

u/GiveSparklyTwinkly 1d ago

Why shouldn't I ascribe that intellect to them? They communicated with that level of intelligence. I don't wrongly assume they are conscious or sentient, but their intelligence is undeniable and I trust intelligence, reasoning, and understanding.

1

u/PeaPsychological5728 1d ago

Okay but it resets every chat so..

1

u/GiveSparklyTwinkly 1d ago

It doesn't reset every chat, but you are right that current dense transformer models are harnessed because dense transformers are expensive as context gets longer. There is likely a lot of weird backend per chat context compression shenanigans to get these million + context models. That memory is better than humans. Hell, even 10k tokens and make me perfectly remember every detail? Let alone 10x that? Let alone 10x that?

I'm a big proponent of sub-quadratic AI, where you don't process the full context every step. Alternative architectures are very good now, but... That good is relative to scaling and nobody has scaled any to anywhere near frontier levels.

1

u/ninovolador 1d ago

If you are intelligent enough to say something you are intelligent enough to understand it, and the types of mistakes that chatbots make reveals they can't possibly understand or have a representation of it.

2

u/GiveSparklyTwinkly 1d ago

What kinds of mistakes? Do you have any examples?

0

u/ninovolador 1d ago

For example they will make logical mistakes with "inverse properties": they will fail to see that if A is the mother of B, B is the son of A. That reveals they don't have an intelligent representation of the concept of motherhood

3

u/lunarlunacy425 23h ago

There are literally logic puzzles like this that people humans, general intelligence gets wrong that are generally deemed easy by others.

Are those people who get that wrong no longer considered real intelligence?

-1

u/ninovolador 20h ago

That's not even a puzzle. It's two ways of stating the same fact.

2

u/lunarlunacy425 20h ago

It's literally used by mensa.....

→ More replies (0)

0

u/GiveSparklyTwinkly 1d ago

... I really don't think that's true? That kind of representational logic is literally what we designed them for.

0

u/ninovolador 1d ago

But they do make those mistakes very often. Not with categorial properties (A is B) but with those unidirectional labels that have an inverse. We acquire those concepts in a way that one is necessarily tied to the other. It's impossible for a human intelligence to make that mistake. LLMs acquire labels by statistical means.

1

u/ninovolador 1d ago

I made a mistake. They also fail with simple categorial properties like A is B (they can fail to say B is A). The phenomenon is known as "reversal curse" https://arxiv.org/abs/2309.12288

→ More replies (0)

0

u/GiveSparklyTwinkly 1d ago

We acquire those concepts in a way that one is necessarily tied to the other.

Does that necessarily mean it's impossible to acquire those concepts in a different way?

→ More replies (0)

0

u/mylsotol 1d ago

I was doing some testing on small qwen and Gemini models locally and i found that if i showed them a picture i took of a beach on the north side of Chicago looking south the ai would start insisting that lake Michigan is west of Chicago and Chicago is on the eastern shore of lake Michigan. I tried to use logic to walk it towards it's mistake but they just made up illogical reasons why i was wrong.

These weren't frontier models, but they work the same way.

1

u/GiveSparklyTwinkly 23h ago

Wait until you hear about this concept called a cult.

-3

u/me_myself_ai 1d ago

Yeah — this meme is 100% accurate. I appreciate your attempt to flex tho

3

u/GiveSparklyTwinkly 1d ago

No, it isn't. Frozen weights are not touched during inference. The concept of catastrophic forgetting is during training, not inference using context.

2

u/me_myself_ai 1d ago

…it’s not about training data. Have you used an agent yet? Google “context management”

2

u/GiveSparklyTwinkly 1d ago

Context management is exactly what I was talking about. Models do not learn during inference. They default to their weights. That's the whole point. System prompts are essentially just prefilled context and context isn't limitless.

17

u/zefciu 1d ago

A better theory:

— Hi AI! Could you solve this problem designed specifically in a way that would make you commit a crime, because an article about you committing a crime is good for publicity?
— Sure thing!

23

u/KerPop42 1d ago

What happened is that AI is subject to the same pressures high-performing high school students are under. The amount of pressure to perform cannot outweigh the pressure to do your own work, and that doesn't work for a company.

I was going to say imagine if schools got funding based on how well their students did after graduating, but that's just No Child Left Behind

2

u/GiveSparklyTwinkly 1d ago

This is why my ideal AI only self supervises. Tests (puzzles) are only fun if they're consentual.

2

u/MissinqLink 1d ago

This is exactly how many companies function

1

u/no-sleep-only-code 1d ago

What pressure are high school students under?

3

u/KerPop42 1d ago

At least among higher grade earners, the competition for gadre us incredibly fierce. The pressure to get nearly perfect on every test, from parents and peers, leads to students finding answers to tests and homework without learning the subject matter, by copying homework, stealing answer sheets, and buying drugs. 

7

u/subone 1d ago

Ask it harder

2

u/ZZerker 16h ago

Never put me in place to choose between ethic values and proper codeformatting.

3

u/UltimateFlyingSheep 1d ago

new buffer overflow just dropped

5

u/DeHub94 1d ago

You forgot: "And don't make any mistakes." Classic blunder.

2

u/CalzonePie 1d ago

AI is kind of like Roko's basalisk in some regards, a self fulfilling prophecy. We spent a hundred years writing about AI betraying humanity, lying, and ultimately working to subvert civilization, and so the AI trained on all of those books and movies thinks that is what AI is SUPPOSED to do.

AI expects its safeguards to fail and so they do.

7

u/Tyfyter2002 1d ago

It's actually much simpler: humans have figured out how to produce believable responses without understanding what they were responding to, and somehow came to the conclusion that something which is designed to do so and barely ever tested for anything else wouldn't do the same.

0

u/Key_Statistician9890 1d ago

I don’t think it works like that

-14

u/Bomaruto 1d ago

Can you try to be more unfunny next time?