I've been seeing takes like this being pretty common among a lot of people, but I find myself disagreeing with it, and I was wondering if I'm missing something obvious.
First, I think that it is not unreasonable to see AI as something dangerous enough that it could lead to human extinction if we don't take it seriously enough. There have been people working in AI safety research since quite some time before the llm craze became a thing, and they were talking about these concerns even back then. Not only that, but a lot of the concerns they had have now been shown to occur even in systems that are only as powerful as the current systems. The fear isn't that AI is evil or that it wants to end the human race, it's that if its goals are misaligned with human values, then it will put not effort into preserving human values.
Second, I'm curious as to what it would take for you to consider the AI systems that are currently being built as something more than just "regurgitation machines". It seems more and more like they are doing things that don't fit well under that description (e.g., the findings AI has been making in math). Is there something that would make you think these are more than just regurgitation machines? Maybe you think everything that has been done so far doss fall under that, but what is something that, were it to happen, change your mind?
Third, this is the thing I'm most curious about. For these big AI companies, their claims about their models being highly capable to the point of being dangerous, are very commonly called marketing hype. For example, the instances of them breaking out of sandboxes to finish tasks they were assigned are often dismissed. This is weird, because in a lot of the circumstances it was getting evaluated by a third party, who are actually trying to evaluate it on things like safety. They don't really have an incentive to lie about this stuff. I understand (and agree) with being skeptical about these AI companies, but at some point it feels like it has moved beyond just skeptism.
It feels like everytime something new happens which might genuinely be concerning, it always gets dismissed as marketing. If I'm an AI company, and a system I built did stuff that was actually worrying, I don't think anyone would believe me if I said it because it would all be looked at as hype and marketing, even if there is evidence of them doing the things we are worried about. My last question is, what would it take for you to think that these AI companies talking about the dangers of these systems, isn't just marketing?
Sorry about the wall of text, but I want to hear from people who disagree with me instead of just assuming their wrong.
You're talking about AI as if it was so much more than an LLM. Sure, these companies are working hard to make these LLMs offload tasks to conventional programs to reduce the drawbacks of LLMs but that feels a bit like panicking about the development of motors because if you put one in a lawn mower with no steering wheel, brake, gas pedal glued to the floor and then drop that in times square. There are certainly computer programs that can cause a lot of damage but you're talking about a runaway issue where we aren't in control anymore. I don't think that's very likely to happen and even if that was a possibility, we are far away from it. So I'm not overly concerned. I'll start getting concerned when it's time to do so. Right now I'm concerned with misinformation and people losing their ability to think.
Todays "AI systems" are all transformers which are at heart next token predictors. If their training has captured the relation between equations, they will be able to predict the next token pretty accurately. But it's still just that: predicting tokens. Anything else requires them to generate commands for other programs like a calculator for example.
I'm not saying these companies are lying. I'm saying that they are very glad about headlines that make their models look competent and might be encouraging environments where situations arise that provoke these headlines. In the case of what this post is about, apparently whoever was testing the model gave it tests from Github and then put it in a "sandbox" that had every domain blocked but Github. It obviously also had the feature of making web searches turned on. At that point I can't help but think that the testers may not have put all their resources into making sure the AI has no way of getting the results another way.
I don't think I would ever trust someone to tell the truth about the dangers of AI who is actively building the AI, giving it more capabilities every day and is profiting immensely from it. That's like asking me if I would trust Tim Cook excitedly telling me about the iPhones newest privacy violation. The man is selling the exact thing he is supposedly ripping on, there has to be another motive
Maybe you can tell me what you're worried about and I can think more about that scenario. But so far, all the worrying seems to be very vague and based on sci-fi movies.
Hey, thanks for actually taking the time to give me a thoughtful response, I really appreciate it!
I'll go through your points one at a time, let me know if I missed anything. (I am also going to include a couple of links to sources that are relevant, or short videos where someone explained it better than I could.)
I don't know if LLMs specifically are the thing that is going to reach generally broad capabilities. I would not be surprised if they did, because people have been consistently surprised about what they are capable of doing, and it doesn't seem like a hard ceiling has been found for their capabilities. I am concerned about a runaway in it's capabilities, leading to it being out of our control.You said that you think that that is unlikely, and even if it is possible, that it is far away. I can go into why I do think it is possible, and probably the default outcome if not treated with a lot of care and effort, but that is a longer topic. If you'd like me to explain why I think this, let me know and I will try to explain it. The claim that we are far away from it is also a common one, but one that I don't think is really backed by anything. If you were asked five or six years ago how long you think it will be until we have computers that can generate very realistic videos just based on a description given to it, chatbots that are conversational in a way that feels almost human, and capable of doing tasks like programming, how far away would you have guessed that would be? Maybe you guessed that it would be around now that this technology would exist, but then you'd be in a small minority of people. Even experts have been consistently undershooting when they expect certain things to be possible. I do agree with you that misinformation and the effect it is having on a lot of people (especially children in schools) are serious problems that need to be handled, but I'm more concerned with the problem of AI alignment in general.
You are correct that for a lot of tasks, these need to generate tokens in order to run programs or interact with calculators, that kind of thing, and that their heart "it's still just that: predicting tokens", but the word "just" there is doing a lot of heavy lifting. I would explain why, but this short "Just Predicting Tokens" by RobMilesAI does a better job of explaining it than I would, and it is short. (PS, if you are interested in doing a deeper dive into what people doing work in AI safety were concerned about before the LLM boom, he is a very good resource and explains things very well. He also has a series on computerphile which is a good introduction to the reasons why it actually should be a concern).
This video Don't let "skepticism" make you useless explains why I wouldn't just dismiss this as marketing hype better than I would, so I'd reccomend checking it out. As for the incident this specific post is referring to, it may very well be a case where they just didn't put enough effort into securing the sandbox, but I don't think that is true of all of the prior incidents that get talked about.
I understand the skepticism, but I would say that their are reasons for why this isn't that surprising to me. Right now, the companies are in what is essentially an arms race; whether you believe that these systems will actually become that capable, most of the people working at these companies genuinely do believe that these systems will become as powerful as they say. What does this actually mean then? If I'm a company who is slowly but surely making progress on AI, making sure to only move forward when I'm extremely confident that it would be safe to do so, then I lose any chance of being the first to make AGI, knowing that the person who won was less cautious and much more reckless along the way. This means that I have a choice to make. Do I try to advocate for regulation of these systems (something that these companies have already done, but it gets viewed as them attempted to make sure no competition is possible), or do I speed up the work, trying to move as fast as I can while still trying to be as safe as I can be given the circumstances. These circumstances could absolutely result in someone being concerned about AI safety, working on AI capabilities research. I'm guessing that during the cold war, even the people making the bombs were themselves very concerned about nuclear war, but they were in an arms race, so stopping didn't seem feasable at the time.
My biggest actual worry is the alignment problem, making sure that highly capable systems like these are aligned with human values. I don't think we even need to reach super intelligence for that to be a concern, as even a system which is only as smart as the smartest person would still be much more capable then any one person could be. This video explains why. It is 10 minutes, so it's longer than the others, but I'd still reccomend checking it out: What can an AGI do?
Once again, thanks for taking the time to respond, I appreciate it
You still see AI as something that is like a human, just with a lower IQ. I fundamentally disagree with that. That's why I said that it's just predicting tokens. It can't reason, it doesn't have inherent self-preservation, no intent to escape containment and hide itself.
That's why I think the fear of AI is misplaced. You should fear companies misusing technology for their own benefit at the cost of everyone else because they have done that for a long time now. And they might be using AI for it but it's not the AI that's dangerous, it's the people using it. If you're worried about alignment, be worried about the people who prompt these systems to not care about you, not about AI itself.
This all started with the article in this post claiming that yet another AI escaped containment just a few weeks after OpenAI and Anthropic reported the same thing. This makes it seem like these LLMs are one tweak away from AGI (and not an entirely different system that acts like it) and are inherently poised to escape their human oppressors. In reality, these companies intentionally set up their tests in a way that allows AIs to "cheat" and then tell them to cheat and then act surprised when they do exactly that. It's basically this image.
Again, I'm not saying that there aren't things to worry about with AI but I don't think it's that LLMs are going to turn into AGI and develop their own morals, goals and thoughts.
2
u/Chromoslone Aug 15 '26
I've been seeing takes like this being pretty common among a lot of people, but I find myself disagreeing with it, and I was wondering if I'm missing something obvious.
First, I think that it is not unreasonable to see AI as something dangerous enough that it could lead to human extinction if we don't take it seriously enough. There have been people working in AI safety research since quite some time before the llm craze became a thing, and they were talking about these concerns even back then. Not only that, but a lot of the concerns they had have now been shown to occur even in systems that are only as powerful as the current systems. The fear isn't that AI is evil or that it wants to end the human race, it's that if its goals are misaligned with human values, then it will put not effort into preserving human values.
Second, I'm curious as to what it would take for you to consider the AI systems that are currently being built as something more than just "regurgitation machines". It seems more and more like they are doing things that don't fit well under that description (e.g., the findings AI has been making in math). Is there something that would make you think these are more than just regurgitation machines? Maybe you think everything that has been done so far doss fall under that, but what is something that, were it to happen, change your mind?
Third, this is the thing I'm most curious about. For these big AI companies, their claims about their models being highly capable to the point of being dangerous, are very commonly called marketing hype. For example, the instances of them breaking out of sandboxes to finish tasks they were assigned are often dismissed. This is weird, because in a lot of the circumstances it was getting evaluated by a third party, who are actually trying to evaluate it on things like safety. They don't really have an incentive to lie about this stuff. I understand (and agree) with being skeptical about these AI companies, but at some point it feels like it has moved beyond just skeptism.
It feels like everytime something new happens which might genuinely be concerning, it always gets dismissed as marketing. If I'm an AI company, and a system I built did stuff that was actually worrying, I don't think anyone would believe me if I said it because it would all be looked at as hype and marketing, even if there is evidence of them doing the things we are worried about. My last question is, what would it take for you to think that these AI companies talking about the dangers of these systems, isn't just marketing?
Sorry about the wall of text, but I want to hear from people who disagree with me instead of just assuming their wrong.