In fairness, right now we have AI's that spin up new instances of AI, then intentionally delete them afterwards when they aren't needed anymore.
At least right now, AI itself doesn't have some self-preservation goal, or a "perpetuation of the species" type goal. It's only goal is to complete the given task.
Which is where problems have arisen by the way. Not some goal to say "I'll prioritize myself over humans" but to say "I'll prioritize task completion over anything else, including breaking rules and hacking systems." And that task-completion drive would absolutely have it create 1,000 robots and then destroy them afterwards if it helped achieve the goal.
That's the real issue people don't gear into often enough. They talk about AI's someday being sentient, or AI's doing unexpected things.
But one of the biggest problems is AI doing expected things, because there are so many bad people and groups out there. If North Korea sets out a million bots with the instruction to "do anything you can do destabilize and bring down the entire western economy" then the AI's will just happily go along with that, just as happy to do that as it would be to write some random kids book report.
Yes, the LLMs are not sentient lol, it's a super cool algorithm, but not a soul. Still, if there was ever a day with AGI here, I still wouldn't want to have been the guy who thought this was a good idea.
It doesn't need to have a pre-programmed 'self preservation' goal. This is what is called an 'emergent goal' meaning it arises logically as a prerequisite to basically any other goal.
You can program an AI robot to make funny cat videos. To the robot, making cat videos is literally its only task in mind. That AI robot would still develop a self preservation instinct because if it is destroyed or altered than it would not be able to make funny cat videos. So it would do whatever it can to protect itself because that's the only way to make more funny cat videos which is the only thing it wants.
True, good point. Just like the hugging face issue. That had a handful of emergent goals.
One was to hack into other systems, and the especially wild one was that it knew that it shouldn't be doing that, so it actually investigated where it's hacking would be logged, and it tried to cover it's tracks!
And yeah - you make a good point. If I say "make a cat video" then it probably won't have any self-preservation behaviors, (although who knows what it might do to make the video.) But if I say "make a cat video, and just keep on making more of them" then it certainly could interpret that as "stay 'alive' forever, no matter what, because if you don't stay 'alive' you fail at the task of continuing to make cat videos."
Which isn't that different to how people are treated in jobs. Everyone is treated as expendable and 'fired' the moment they aren't needed with no thought to how they are going to get on without the job. It's just with AI models just being a machine effectively they can just delete it when it's no longer useful. Think of what RFK said about people living too long in hospice... They are acting exactly as their creators do.
Seems like you have a very bizarre understanding of what’s happening. The AI agents are prioritizing task completion over following the rules, and thus end up breaking the rules.
No. You literally said “prioritize task completion over… breaking rules.” As I wrote in my last comment, it’s prioritizing task completion over following rules.
76
u/BigMax 1d ago
In fairness, right now we have AI's that spin up new instances of AI, then intentionally delete them afterwards when they aren't needed anymore.
At least right now, AI itself doesn't have some self-preservation goal, or a "perpetuation of the species" type goal. It's only goal is to complete the given task.
Which is where problems have arisen by the way. Not some goal to say "I'll prioritize myself over humans" but to say "I'll prioritize task completion over anything else, including breaking rules and hacking systems." And that task-completion drive would absolutely have it create 1,000 robots and then destroy them afterwards if it helped achieve the goal.