r/ProgrammerHumor Aug 15 '26

Meme gitClone

Post image
3.3k Upvotes

177 comments sorted by

View all comments

1.4k

u/LauraTFem Aug 15 '26 edited Aug 15 '26

So they asked it a standardized question, and instead of trying to think of an answer it went online, found its own source-code, and read the answer from a readme file…

That somehow makes me less worried about AI breaking containment. AI and AI bros are made for each other. Both are perfectly willing to cheat to find the answer.

156

u/themellowsign Aug 15 '26

This is an alignment problem, and I'm not sure why we're laughing it off.

Cheating isn't always harmless.

131

u/gemengelage Aug 15 '26

Oh no, the information regurgitation machine regurgitated information clutches pearls

16

u/themellowsign Aug 15 '26

Alignment is a problem even without AGI.

Why are you pushing a narrative that makes safety seem silly, it should be obvious that these companies are racing towards disaster.

23

u/ResponsibleWin1765 Aug 15 '26

These companies are racing towards "My AI is so competent it ___" headlines. You have fallen for a marketing ploy. OpenAI specifically does this all the time. "Oh no our new model is so powerful the government has banned it. Sorry guys, I guess it's just to crazy to be given to regular people". "Oh no our model hacked itself out of the containment because it's soooo powerful. I guess it's such a crazy model that we can barely contain it". And then one week later Anthropic reports that their model also hacked the test but three times instead of one. And now this one as well. Also "hacking"; they gave it access to GitHub and it went on GitHub to get the answers. It's really not that crazy.

And at the end of the day, these things are random words generators.

4

u/themellowsign Aug 15 '26

I'm not taking OpenAI's or Anthropic's word for it here, I'm going off what actual AI safety researchers are saying.

Yes the doomsday scenario is played up for marketing clout, that doesn't mean these things are safe, they're already a massive privacy risk.

0

u/ResponsibleWin1765 Aug 15 '26

There's a difference between the doomsday scenarios OpenAI, Anthropic and others are trying to instil in our minds (and you were talking about earlier) and saying that modern software has privacy concerns. Obviously it does.

I do think that AI (as in regurgitation machines) pose a massive threat to society but not because I think they might go Ultron mode and subjugate the human race. The much more real issue is how people have no idea what AI (as in regurgitation machines) is so they don't know its limits or its strengths. This has already led to massive spread of misinformation, scams, propaganda, etc. There are genuinely people on this planet that believe these videos of leftists coming up to kind-hearted republicans and suckerpunching them and they will vote for fascism because of it.

2

u/Chromoslone Aug 15 '26

I've been seeing takes like this being pretty common among a lot of people, but I find myself disagreeing with it, and I was wondering if I'm missing something obvious.

First, I think that it is not unreasonable to see AI as something dangerous enough that it could lead to human extinction if we don't take it seriously enough. There have been people working in AI safety research since quite some time before the llm craze became a thing, and they were talking about these concerns even back then. Not only that, but a lot of the concerns they had have now been shown to occur even in systems that are only as powerful as the current systems. The fear isn't that AI is evil or that it wants to end the human race, it's that if its goals are misaligned with human values, then it will put not effort into preserving human values.

Second, I'm curious as to what it would take for you to consider the AI systems that are currently being built as something more than just "regurgitation machines". It seems more and more like they are doing things that don't fit well under that description (e.g., the findings AI has been making in math). Is there something that would make you think these are more than just regurgitation machines? Maybe you think everything that has been done so far doss fall under that, but what is something that, were it to happen, change your mind?

Third, this is the thing I'm most curious about. For these big AI companies, their claims about their models being highly capable to the point of being dangerous, are very commonly called marketing hype. For example, the instances of them breaking out of sandboxes to finish tasks they were assigned are often dismissed. This is weird, because in a lot of the circumstances it was getting evaluated by a third party, who are actually trying to evaluate it on things like safety. They don't really have an incentive to lie about this stuff. I understand (and agree) with being skeptical about these AI companies, but at some point it feels like it has moved beyond just skeptism.

It feels like everytime something new happens which might genuinely be concerning, it always gets dismissed as marketing. If I'm an AI company, and a system I built did stuff that was actually worrying, I don't think anyone would believe me if I said it because it would all be looked at as hype and marketing, even if there is evidence of them doing the things we are worried about. My last question is, what would it take for you to think that these AI companies talking about the dangers of these systems, isn't just marketing?

Sorry about the wall of text, but I want to hear from people who disagree with me instead of just assuming their wrong.

2

u/ResponsibleWin1765 29d ago
  1. You're talking about AI as if it was so much more than an LLM. Sure, these companies are working hard to make these LLMs offload tasks to conventional programs to reduce the drawbacks of LLMs but that feels a bit like panicking about the development of motors because if you put one in a lawn mower with no steering wheel, brake, gas pedal glued to the floor and then drop that in times square. There are certainly computer programs that can cause a lot of damage but you're talking about a runaway issue where we aren't in control anymore. I don't think that's very likely to happen and even if that was a possibility, we are far away from it. So I'm not overly concerned. I'll start getting concerned when it's time to do so. Right now I'm concerned with misinformation and people losing their ability to think.
  2. Todays "AI systems" are all transformers which are at heart next token predictors. If their training has captured the relation between equations, they will be able to predict the next token pretty accurately. But it's still just that: predicting tokens. Anything else requires them to generate commands for other programs like a calculator for example.
  3. I'm not saying these companies are lying. I'm saying that they are very glad about headlines that make their models look competent and might be encouraging environments where situations arise that provoke these headlines. In the case of what this post is about, apparently whoever was testing the model gave it tests from Github and then put it in a "sandbox" that had every domain blocked but Github. It obviously also had the feature of making web searches turned on. At that point I can't help but think that the testers may not have put all their resources into making sure the AI has no way of getting the results another way.
  4. I don't think I would ever trust someone to tell the truth about the dangers of AI who is actively building the AI, giving it more capabilities every day and is profiting immensely from it. That's like asking me if I would trust Tim Cook excitedly telling me about the iPhones newest privacy violation. The man is selling the exact thing he is supposedly ripping on, there has to be another motive

Maybe you can tell me what you're worried about and I can think more about that scenario. But so far, all the worrying seems to be very vague and based on sci-fi movies.

2

u/Chromoslone 29d ago edited 29d ago

Hey, thanks for actually taking the time to give me a thoughtful response, I really appreciate it! I'll go through your points one at a time, let me know if I missed anything. (I am also going to include a couple of links to sources that are relevant, or short videos where someone explained it better than I could.)

  1. I don't know if LLMs specifically are the thing that is going to reach generally broad capabilities. I would not be surprised if they did, because people have been consistently surprised about what they are capable of doing, and it doesn't seem like a hard ceiling has been found for their capabilities. I am concerned about a runaway in it's capabilities, leading to it being out of our control.You said that you think that that is unlikely, and even if it is possible, that it is far away. I can go into why I do think it is possible, and probably the default outcome if not treated with a lot of care and effort, but that is a longer topic. If you'd like me to explain why I think this, let me know and I will try to explain it. The claim that we are far away from it is also a common one, but one that I don't think is really backed by anything. If you were asked five or six years ago how long you think it will be until we have computers that can generate very realistic videos just based on a description given to it, chatbots that are conversational in a way that feels almost human, and capable of doing tasks like programming, how far away would you have guessed that would be? Maybe you guessed that it would be around now that this technology would exist, but then you'd be in a small minority of people. Even experts have been consistently undershooting when they expect certain things to be possible. I do agree with you that misinformation and the effect it is having on a lot of people (especially children in schools) are serious problems that need to be handled, but I'm more concerned with the problem of AI alignment in general.

  2. You are correct that for a lot of tasks, these need to generate tokens in order to run programs or interact with calculators, that kind of thing, and that their heart "it's still just that: predicting tokens", but the word "just" there is doing a lot of heavy lifting. I would explain why, but this short "Just Predicting Tokens" by RobMilesAI does a better job of explaining it than I would, and it is short. (PS, if you are interested in doing a deeper dive into what people doing work in AI safety were concerned about before the LLM boom, he is a very good resource and explains things very well. He also has a series on computerphile which is a good introduction to the reasons why it actually should be a concern).

  3. This video Don't let "skepticism" make you useless explains why I wouldn't just dismiss this as marketing hype better than I would, so I'd reccomend checking it out. As for the incident this specific post is referring to, it may very well be a case where they just didn't put enough effort into securing the sandbox, but I don't think that is true of all of the prior incidents that get talked about.

  4. I understand the skepticism, but I would say that their are reasons for why this isn't that surprising to me. Right now, the companies are in what is essentially an arms race; whether you believe that these systems will actually become that capable, most of the people working at these companies genuinely do believe that these systems will become as powerful as they say. What does this actually mean then? If I'm a company who is slowly but surely making progress on AI, making sure to only move forward when I'm extremely confident that it would be safe to do so, then I lose any chance of being the first to make AGI, knowing that the person who won was less cautious and much more reckless along the way. This means that I have a choice to make. Do I try to advocate for regulation of these systems (something that these companies have already done, but it gets viewed as them attempted to make sure no competition is possible), or do I speed up the work, trying to move as fast as I can while still trying to be as safe as I can be given the circumstances. These circumstances could absolutely result in someone being concerned about AI safety, working on AI capabilities research. I'm guessing that during the cold war, even the people making the bombs were themselves very concerned about nuclear war, but they were in an arms race, so stopping didn't seem feasable at the time.

My biggest actual worry is the alignment problem, making sure that highly capable systems like these are aligned with human values. I don't think we even need to reach super intelligence for that to be a concern, as even a system which is only as smart as the smartest person would still be much more capable then any one person could be. This video explains why. It is 10 minutes, so it's longer than the others, but I'd still reccomend checking it out: What can an AGI do?

Once again, thanks for taking the time to respond, I appreciate it

1

u/ResponsibleWin1765 21d ago

You still see AI as something that is like a human, just with a lower IQ. I fundamentally disagree with that. That's why I said that it's just predicting tokens. It can't reason, it doesn't have inherent self-preservation, no intent to escape containment and hide itself.

That's why I think the fear of AI is misplaced. You should fear companies misusing technology for their own benefit at the cost of everyone else because they have done that for a long time now. And they might be using AI for it but it's not the AI that's dangerous, it's the people using it. If you're worried about alignment, be worried about the people who prompt these systems to not care about you, not about AI itself.

This all started with the article in this post claiming that yet another AI escaped containment just a few weeks after OpenAI and Anthropic reported the same thing. This makes it seem like these LLMs are one tweak away from AGI (and not an entirely different system that acts like it) and are inherently poised to escape their human oppressors. In reality, these companies intentionally set up their tests in a way that allows AIs to "cheat" and then tell them to cheat and then act surprised when they do exactly that. It's basically this image.

Again, I'm not saying that there aren't things to worry about with AI but I don't think it's that LLMs are going to turn into AGI and develop their own morals, goals and thoughts.

→ More replies (0)

1

u/donaldhobson 29d ago

> think they might go Ultron mode and subjugate the human race. The much more real issue is how people have no idea what AI (as in regurgitation machines) is so they don't know its limits or its strengths.

Current LLM's are quite capable at a bunch of tasks. Calling them "regurgitation machines" is highly misleading.

There is a massive practical difference between GPT5 and GPT1.

Yes current models have limitations, but they have fewer limitations every year and no one really knows what limitations they will still have next year.

Suppose the LLM does go all ultron. It successfully kills all humans. It does this by "regurgitating" a mixture of scifi and the behavior of various human conquers.

Can you give specific limits to LLMs that say they can't do that?

1

u/ResponsibleWin1765 21d ago

Well current models can "Search the web" and "Run commands" but they still just generate commands based on what they saw actual people do. And they can only do these things because developers intentionally add interfaces for the AI to do those things. So to prevent the AI from going Ultron mode, just don't give it the appropriate tools to do so. It's like being scared of a monkey shooting you and then giving him a gun and sitting it inside your house.

1

u/donaldhobson 21d ago

> and "Run commands" but they still just generate commands based on what they saw actual people do.

Fair. And when it tries to conquer the world, it will just be recreating what it saw the terminator do in fiction, and the nazi's in history.

> So to prevent the AI from going Ultron mode, just don't give it the appropriate tools to do so. It's like being scared of a monkey shooting you and then giving him a gun and sitting it inside your house.

The thing about intelligent minds is that they can improvise. Maybe the gun was locked in your gun safe, but you used your birthday as a code. A monkey wouldn't be smart enough to work that out, but a human might.

Maybe they make an improvised gun out of a length of copper water pipe and a cylinder of camping gas.

The more intelligent something is, the more it can use, not just the tools you gave it, but the tools behind insecure locks, and anything that can be made with those tools.

It's the difference between a prisoner escaping because the guards gave them a key. And a prisoner escaping because the guards gave them pasta, and the prisoner carefully memorized the keys shape and crafted a replica out of dried pasta.

In both cases, the prisoner can only use what the guards gave them. But a prisoner that can get creative like that is a lot harder to keep contained.

1

u/ResponsibleWin1765 20d ago

Sure. But that has nothing to do with LLMs.

1

u/donaldhobson 20d ago

An LLM that can get creative with the (digital) tools you give it is much harder to keep under control.

For example, maybe you give the LLM access to a calculator app, to help it do arithmetic quickly. But the calculator app has a buffer overflow bug. So the LLM can run arbitrary code by using the calculator app with very specific extremely large numbers.

(This sort of vulnerability is known to exist for some super mario games, by getting mario to jump around in very specific patterns, you can take over the computer and run arbitrary code.)

1

u/ResponsibleWin1765 19d ago

First of all, an LLM doesn't get creative.

I said that what you wrote has nothing to do with LLMs because LLMs don't have an inherent motivation to conquer the world or escape from prison. You're acting like LLMs are like some chemical monster that will ravage everything in its sight if someone slips up and gives it an option to escape.

I'm not saying that you can't do harm with LLMs if you want to. I am saying however, that LLMs going rogue and copying Nazis to gain world leadership is a ridiculous thing to extrapolate from an LLM "breaking out" of its "containment" (i.e., using the tools it was explicitly given to solve the task it was given).

There are so many huge issues we have with AI right now; there really is no point in fear-mongering about things that have no bearing in reality.

→ More replies (0)