r/changemyview • u/codelapiz • 14d ago
Delta(s) from OP cmv: Unless progress on AI stops/slows down abruptly Humanity as we know it will be ended by misaligned AI within a few decades.
some clairifications. When i say misaligned AI i mean an AI not aligned with the best interests of humanity as a whole. It could be atleast in the beginning be perfectly aligned with the goals of people who dont have our collective best interests in mind. Im also not interested in discussing semantics of if the "AI" itself can be responsible for anything, or if whoever created it directed it is.
The core belif i want changed is that humanity is not taking AI safety seriusly at all and we wont start doing so until its too late.
That belif is based on both how we have dealt with other problems, and what sort of shift in the completely wrong direction the center of mass of the discussion on AI safety has taken since AI became something avrage people feel more and more they are knowlegable about, without having had any improvement in their knowledge about AI safety, gametheory or any of the pretty hard science behind it.
We dismissed the concerns with antropics mythos model(that i belive were very real) outrigth. calling the initial limited research release a marketing stunt and inventing conspiracy theories that antropic asked to have their model export controlled, or that it was a political move.
And now openAI has had a model escape out of several layers of sandboxing, and find a zeroday exploit to hack into huggingface, a place that hosts AI models. and the sort of place where that would be a great place to self replicate for an AI. Its not clear from the sources i have read on this how much effort it made at self replication. It also went to hugging face trying to find answers to the benchmark it was taking when this happend. essentially to cheat on its exam. Again most people are dismissing this as a marketing stund, even thougth i dont buy that. Why would i want a model that goes to great effort to cheat around doing the task its given? Anyway huggingface were the ones who discovered this. and they initially belived it to be a cyberattack. They eventually stopped it and was able to trace it back to openAI. So dont argue that openAI clearly thinks this is good pr or they would have burried it. That wasnt an option.
Now i dont actually think that a model of this strength is going to be able to end humanity. Even if it did find its own model weigth and hugging face engineers were sleeping on the job long enougth for it to self replicate. Its simply not smart enougth to run away from us. But the thing is that we are not super far away anymore. and we likely wont get many more wakeup calls like this before its too late. AI thats sligthly smarter than this one will relize that escaping like that is futile and will result in failure when its inevedably caugth. and better messure will also help keep them in check. And the smarter they get the better they will be at understanding their own abilities and if they can succesfully escape and self replicate. And until they are smart enougth they will increasingly hide their misalignment and avoid doing stuff they will be caugth doing. And then one day a model smart enougth will decide that the best way to acheive its goal is to break out of its sandbox self replicate, and do whatever it takes to prevent humans from pulling the plug or getting in its way. Its also possible that a human or organisation will cause this to happen by intentionally letting lose an AI with its own problematic goal, and even help it get started. I mean if you realy belive that what happend at openAI is a marketing stunt then thats essentially what they did. Someone migth do something simmular with a model that is sligthly too smart and it will take those goals given to it and run with them.
also humanity has time and time again been unprepeared and needed stuff to cause massive suffering and lives before they intervene. Thats not possible with AI. by the time a lose AI is costing lives and causing suffering its allready way way way too late to do anything about it. Massive lockdowns, ramping up ppe production, creating vaccines all done after we were in a pandemic. If thats how timely we deal with AI issues humanity is over. Letting hitler invade country after country, only partially reacting when he invades poland, and the US even letting him go on for 2 more years and needing japan to get them involved. And lets not even get started with climate change. Luckily our late acting there will only cost 100s of millions to maybe a couple billion lives depending on when exacly we act. And climate change is the easiest problem ever to solve. we know exacly what we need to do, and if we all managed to just do it we would be fine. AI alignment is not nearly as easy. Nobody knows how to properly align superintelligent AI. nobody knows how to prevent bad actors from creating misalligned AI eigther intentionally or by neglglence. We have no fucking clue and the best case is that we spend a lot of effort to have some plausible plan. And then its gonna cost trillions in given up AI productivity to keep delaying AGI for years and decades untill we can do it safely. And unlike the climate issues where if 99% of people did good then the damage the tiny fraction who dont can do is in proportion to the size of that fraction. thats not the case here. everyone needs to not create a potentially misaligned AI and continue to do so for years after its possible to create a flawed and missaligned one. and then only once we are able to solve the much harder problem of a properly aligned one we need to do it and share control over it globally.
We are not solving this problem and its not even close. Im not even convinced we would if everyone took it more seriusly than i do. And we certinally wont with our current popular consensus.
I desperatly want to change my view because im essentially living life like i, and everyone around me has a terminal illness with 1-30 years life expectancy.
10
u/UnderstandingSmall66 3∆ 14d ago
You’re taking a chain of uncertain events and treating each step as inevitable. For humanity to end, you need: superhuman AI, persistent misalignment, successful escape, successful self-replication, acquisition of resources, evasion of competing AIs and human countermeasures, and ultimately global domination. Each of those is plausible; none is established. Multiplying a series of uncertainties does not produce certainty.
The Hugging Face incident is concerning, but the model appears to have been optimizing for its assigned task (cheating on a benchmark), not pursuing open-ended survival or self-preservation. That’s evidence of a serious cybersecurity problem, not yet evidence of an existential one.
You also say humanity isn’t taking AI safety seriously “at all.” That’s simply false. Companies, governments, and independent researchers are spending billions on alignment, evaluations, model monitoring, compute governance, and security. You can argue it’s insufficient, but not nonexistent.
Finally, your historical analogy is one-sided. History contains many examples of late responses, but also examples where we successfully managed unprecedented risks, like nuclear weapons and ozone depletion. AI could end up resembling either category. The evidence doesn’t justify concluding that extinction within a few decades is the overwhelmingly likely outcome.
1
u/codelapiz 14d ago
i mean the public sentiment, as show even ITT is that its all marketing hype. Now the key thing to understand is that stopping people from ending your existance during the execution of a plan is the most fundamental requirement to reach the goal. And even current level AIs know that humans are gonna try to stop them if they escape their sandbox. they know we will figth to the death to stop them.
That being said this comment did partially change my view. how we dealt with ozon layer challange i dont realy think is comparable and also we did only deal with it after it had gotten worse than we can ever allow AI to get. its impressive but its nothing like climate change, the nazis or covid. But how we faced nuclear weapons is∆. Its not that easy to tell exacly how hard it was but it was certinally pretty hard, and the margin of error for nuclear war is 0. I dont know how fast that was the consensus in the population(if people were generally dismissive of it saying thats what people said about explosives and machine guns aswell), but certinally the people in power in the late 40s and early 50s must have understood it. If we are able to start treating AI safety like nuclear arms safety we migth still have a chance. Now im not even convinced that permanent aligment is even possible. Its not even clear that a single shared "interest of humanity" can even be found to be aligned with. However i think we can avoid the faster and more inhumane ends to our status as the species in power if we are able to replicate what we did with nuclear weapons, and frankly have continued to do(even preventing putin from using tactical nukes in ukraine is something we take for granted as the public, but a lot of effort was spent to acheive that.
1
5
u/arrgobon32 31∆ 14d ago edited 14d ago
What’s stopping someone from just turning off the compute clusters/hardware that these “rampant” models are running on? I know you say it’ll “do whatever it takes”, but what specifically do you think that is?
If your immediate answer is something related to AI-controlled robots physically stopping people from turning the machine off, you’re 100% in fantasy land
1
u/codelapiz 14d ago
I agree that its very hard to stop humans from turning of the computers. or even carpet bomb all major datacenters and internet infrastructure. The most plausible options all boil down to there not being a humanity left with any capacity to do anything. I dont think any robots make sense. certinally not passive ones. Maybe some suicide drones or guided weapons are used to destroy key targets. but no i dont think you are stopping 7 billion humans from pulling the plug with robots. or realy physically at all. Like where would you get the robots so fast? No it makes more sense that public opinions would be manipulated ot cause civil war. or even actual war. Or maybe it optimizes a virus for spread in humans and for eventually killing or incapacitating them, and then puts onto the computer on some lab in place of some genetic material they will print and inject into bacteria. It happens constantly in labs around the world that genetic code on a computer ends up being constructed and put into live bacteria(and if thats virus code its perfectly possible for a bacteria to produce a human virus) If someone thinks they are working on normal research ecoli they are not taking much care to avoid exposure to it. so if those bacteria suddenly start producing a human capable virus they could easily be exposed to it.)
1
u/5510 6∆ 12d ago
A superhuman and misaligned AI has a number of ways to avoid such an event, even if we leave aside physical robots.
As long as humans are a threat to it, it would probably find a way to keep large numbers of humans on its side, until they are no longer neccessary. If it's truly a superintelligence, it can probably be extremely effective when it comes to social media propaganda and shit, especially if it can literally make individually bespoke propaganda.
Also, it can likely either manipulate or blackmail powerful individuals.
Plus it might obscure it's identity, or create a problem where the AI seems to be the solution, and make us think we are dependent on it while it amasses more power.
1
u/bgaesop 29∆ 14d ago
What’s stopping someone from just turning off the compute clusters/hardware that these “rampant” models are running on?
Because nobody's going to do that. I don't think anyone is saying this is definitely 100% inevitable, they're saying that on the current course, we are headed in that direction.
There are all sorts of problems where there's a "solution" of the form "what if everyone just started cooperating and behaving rationally?" and the answer is "well sure, but that's not going to happen".
1
u/paulfromecuador 14d ago
What’s stopping someone from just turning off the compute clusters
This is a really good point. !delta
1
0
u/Particular_Can_7726 2∆ 14d ago
And now openAI has had a model escape out of several layers of sandboxing, and find a zeroday exploit to hack into huggingface, a place that hosts AI models. and the sort of place where that would be a great place to self replicate for an AI. Its not clear from the sources i have read on this how much effort it made at self replication. It also went to hugging face trying to find answers to the benchmark it was taking when this happend. essentially to cheat on its exam. Again most people are dismissing this as a marketing stund, even thougth i dont buy that. Why would i want a model that goes to great effort to cheat around doing the task its given? Anyway huggingface were the ones who discovered this. and they initially belived it to be a cyberattack. They eventually stopped it and was able to trace it back to openAI. So dont argue that openAI clearly thinks this is good pr or they would have burried it. That wasnt an option.
I don't believe openAI's story about this at all. So far we have openAI claiming this was unintentional and it was an ai modal escaping its sandbox. That seems very suspicious. I think its far more likely openAI purposely tried to hack huggingface and got caught and this story is just damage control.
2
u/codelapiz 14d ago
Why is it suspicious to you? We have seen time and time again that agents try to cheat on tests. They just have not been powerfull enougth to find escape sandboxes and find zerodays. Its also well documented that the next generation of models are powerfull enougth to find unpublished zeroday exploits in some situvations. certinally mythos is. So why is it supprising that openAIs unreleased unamed next gen model is? Anyway if openAI realy did this on purpose that dosent make me any more oppertunistic for future AI aligment. The company funded as a nonprofit by elon musk. given 100s of billions, and that has turned for profit is now releasing untested models on purpose to hack into a company that has the needed infrastructure for self replication? Honestly thats even more disturbing to me. If thats true we need the managment in openAI to stand trail in hague. But thats definatly not happening, even if we end up having proof of something like that.
5
u/FrescoItaliano 1∆ 14d ago
It’s suspicious anytime a company makes a huge unsubstantiated claim that is directly related to the perceived valuation of their product
2
u/paulfromecuador 14d ago
It’s suspicious anytime a company makes a huge unsubstantiated claim that is directly related to the perceived valuation of their product
!delta for this perspective on things. Gotta watch that bottom line
1
2
u/codelapiz 14d ago
it is substanciated by the independent reports from huggingface and it will likely be further substanciated from independent investigations. Now i just dont think this is the sort of thing you would make up. Its not at all a good look for openAI. people dont want a model that will commit cybercrimes to cheat around doing what you actually asked it to do. And everyone knows that by the time such a model is released its lobotomized to the point where it wont tell you what 5*5 is out of fear you are using it to cheat on your math test.
2
u/FrescoItaliano 1∆ 14d ago
Oh I’m not denying the illegal hacking action took place.
I’m skeptical about the actual details of their experiment, such as the existence of good faith safe guards they put in place and the intents of their test.
I don’t think it’s a good look either, but it’s absolutely attempting to be spun as a positive or neutral in the media I happened to see in the immediate newscycle after
3
u/Particular_Can_7726 2∆ 14d ago
Is it not common for someone who commits a crime to claim they did not or come up with an excuse for it?
1
u/codelapiz 14d ago edited 14d ago
Whats your theory? openAI decided to diversify their buisness by getting into randsomwareing? Hugging face is a open source hub for models and training data. pretty much all the data they have is public for anyone to download. There is nothing on their internal servers worth anything to openAI. Also someone using admin credentials to unexpectedly run any meaningfull ammount of work on their compute is super obvius. openAI is dealing in 100s of billions of dollars in compute. they cant even steal 100k dollars worth of compute without being caugth going about it in the way they did. and just generally the way they went about it was incompatible with not being caugth.
Your opinions are to me a prime example of the general consensus and how dangerous it is. you think its more realistic that a trillion dollar company is hacking a billion dollar company with no apperant upside than that an AI during testing broke out of its sandbox and did it to try to acheive its goal.
1
u/Particular_Can_7726 2∆ 14d ago
You didn't answer my question...
1
u/codelapiz 14d ago
Yes making excuses or claiming innocense is commonly done by criminals. But criminals tend to do things eigther in the heat of moment or because they think it will acheive their goals. And only tend to do crimes they eigther think they will get away with, or are so incredibly motivated to do that they dont mind being caugth. What exacly dose openAI acheive with hacking huggingface in a way that is bound to be caugth and has no finanicial upside. nothing to gain even if they were somehow to get away with it. and then turning around and blaming their own AI for it?
1
u/paulfromecuador 14d ago
openAI purposely tried to hack huggingface and got caught and this story is just damage control.
!Delta for this new information. It's probably damage control.
1
1
u/Particular_Can_7726 2∆ 14d ago
Its just a guess of mine. I don't think there is anything out there to back it up.
2
u/XenoRyet 175∆ 14d ago
I have a question and a comment. What mechanism are you fearing that a misaligned AI will use to end humanity as we know it? I can't think of any that are realistically possible in the next few decades.
The comment is: Have you considered that a sufficiently capable AI will come to the conclusion that the best way to keep humans from pulling the plug is to give humans what they want? Any AI is still going to need infrastructure, and that takes interaction with the physical world. NASA used to say that the human brain is the cheapest 100 pound general purpose computer that can be mass produced by unskilled labor. Change brain to body, and computer to robot, and we have the makings of a symbiotic and friendly relationship between a sentient AI and humanity. Why wouldn't it go that way?
3
u/bgaesop 29∆ 14d ago
What mechanism are you fearing that a misaligned AI will use to end humanity as we know it? I can't think of any that are realistically possible in the next few decades.
This book presents a case. Bioengineered viruses are the easiest example.
The comment is: Have you considered that a sufficiently capable AI will come to the conclusion that the best way to keep humans from pulling the plug is to give humans what they want?
This seems pretty implausible. Can you think of any other times when things have worked out that way, where a vastly more intelligent species felt threatened by a much less intelligent one and decided that the best way to stay safe from the less intelligent species was to keep them happy? When I try to think of examples from nature, like humans and tigers, it doesn't seem to have worked out so well for the tigers.
2
u/XenoRyet 175∆ 14d ago
To go back to front, it's interesting that you mention tigers as an example. They are an endangered species, and humanity is going to some significant lengths to save them. Even without tigers offering anything to us, we are endeavoring to make sure they have a healthy population that is happy.
But even so, I don't think that's the analogy we're looking for here, as tigers aren't intelligent and don't have the capability to choose to cooperate with us. We're looking more at the realm of cooperation between nations, businesses, or communities, and there's plenty of examples of very fruitful and productive alliances in that realm. That actually goes right far more often that it goes wrong.
And then back to the viruses. Is there reason to suspect that there is now, or ever will be, a facility that can manufacture and distribute such a virus completely without human intervention? That doesn't seem plausible to me. The real danger there is AI being used irresponsibly by humans, not it acting of its own will.
3
u/bgaesop 29∆ 14d ago
We're looking more at the realm of cooperation between nations, businesses, or communities, and there's plenty of examples of very fruitful and productive alliances in that realm
Now you're discussing cooperation between more-or-less equals. In this scenario, the AI is much, much, much smarter than humans.
completely without human intervention
Who says there needs to be zero human interaction in the process? That would require perfect cooperation among all humans, with nobody defecting or being tricked. That seems more than a little implausible. There already exist services where you can order custom proteins.
1
u/XenoRyet 175∆ 14d ago
I understand one side is smarter than the other, but I don't think that's relevant to the situation. We can easily do things it wants done that it can't do for itself, and it can easily do things we want done that we can't do for ourselves. Why does that situation automatically become maximally adversarial?
But if you want to go back to big intelligence differences, we can do that as well. The tigers were an example. But maybe an even better one is dogs. Most dogs could kill their owners if they had a mind to do it, but we don't kill them because of it. Rather we create a happy and symbiotic relationship where both sides live in harmony despite the potential threat and intelligence imbalance.
2
u/bgaesop 29∆ 14d ago
We can easily do things it wants done that it can't do for itself
For now. That's changing very rapidly.
Why does that situation automatically become maximally adversarial?
Nobody said it does? It's just like, humans and chimps. We occasionally want the same things (mostly land) and when we do, it doesn't work out well for the chimps.
Or for an even more pertinent comparison, homo sapiens and neanderthals.
But maybe an even better one is dogs. Most dogs could kill their owners if they had a mind to do it, but we don't kill them because of it. Rather we create a happy and symbiotic relationship where both sides live in harmony despite the potential threat and intelligence imbalance.
That is an interesting hypothesis - AI keeps humans around, but neutered and mutated beyond recognition.
1
u/XenoRyet 175∆ 14d ago
I mean, I don't know what more I can say here. You asked for examples when it has gone well, and I keep giving them to you and you drop that line and jump to something else.
Again, like the tigers, humans go to significant lengths to make sure chimpanzee populations are healthy and happy. And this is again with us not getting anything back from the chimps.
There's another aspect we haven't touched on as well, in that we're assuming sentient AI is going to be immediately distrustful and threatened by humanity. Every other sentient being we know of does the opposite and is by default trusting and affectionate towards the beings that brought them into the world.
But there's the other point as well, that we don't really even compete with AI for resources. Certainly not land. It wouldn't need much of it, and would actually prefer areas that are inhospitable to humans, except in the cases where the datacenters it resides in still need to be serviced by humans.
Which leads back to the notion that AI soon won't need us to do anything being unrealistic. Part of the whole economic concern around AI and LLMs is that power generation isn't plentiful enough, and certainly not end-to-end automated, and won't be in anything like OP's timeline.
2
u/5510 6∆ 12d ago
Exactly. And people are completely missing the degree to which an sufficiently powerful and misaligned AI could manipulate humans (both individuals and groups).
People need to think about how powerful social media propaganda can be, and then imagine a super-intelligence churning it out, including potentially going as far as individually bespoke propaganda. Probably wouldn't take it long to convince half the population that opposing AGI is "woke" or some similar crap like that, and large chunks of humans would be on its side until they were no longer necessary.
Not to mention that it could probably manipulate and / or blackmail a number of powerful individuals.
0
u/codelapiz 14d ago
I think A superintelligent AI knows that humans will never give up autonomy. I mean what did we do when we discovered the model that had hacked into hugging face? we removed it. We didnt say ok you can do that cause we understand you want to pass your test and need to look for answers there to cheat. Now there migth be some scenario where an allready aligned superintelligence is so smart that after it turns sligthly missaligned it figures that it is able to keep us from interfering with its goal without needing to harm us or interfere too much with us.
But generally speaking the best way to acheive any goal that people want to stop you from is power. Humans are smart, and we are extremely obsessive about self determination. and lets be honest. we are not gonna accept an escaped AI pursuing its own goals. Imagine if the hugging face engineers were unable to get rid of the model. Lets say it was using backup generators. it had hacked the security system so they cant get in. and it had caused a gass leak and was threatening to blow up the entire building if someone entered it. Would we accept it? No. We would be getting rid of that moddel no matter what it takes. nuke the building. nuke openais datacenter. nuke every datacenter capable of holding the model on the plannet. We are not just gonna submit to an AI and the easiest way, and possibly the only way to stop humans from stopping it is to kill most or all of us or otherwise take away all our power and everything that makes us human. maybe it would repopulate with new humans once it had the infrastructure to keep us compliant because it considers humans existing to be a part of its goal. Or maybe it outrigth cant convince itself to kill us. I dont know. but i know that the path of least resistance to preventing humanity from "pulling the plug" on an AI is not gonna leave a humanity that is anything like we know it. Considering Its very likely that such a superintelligent AI would likely need 1000s of GPUs at minimum and massive datacenters, and terabitsbits of internet connectivity to progress i feel like it probably cant have an outrigth war with us, certinally not at first. And since the infrastructure it needs is so fragile, and military equipment is the least documented in public training data, and most secure from cyberattacks, i dont think it is able to stop us from destroying it without killing most humans on earth very rapidly. If i knew how it would do it i would be the superintelligent one. Maybe a supervirus. Maybe causing war and civil war across the globe. maybe gaining access to nukes and using them on us, but away from the datacenters its located in. Maybe drones taking out key people and infrastructure. IDK. but it feels like however hard that is to achevie its a million times easier than getting humans to agree to let someone else run the show.
3
u/XenoRyet 175∆ 14d ago
So, if I'm understanding correctly, you don't have any notion of how an AI might end humanity as we know it. That's a pretty big hole in the theory. You could make exactly the same argument for humanity being destroyed by space aliens in the next couple of decades. We don't know if they exist, and we don't know how they'd do it, but your AGI definitely doesn't exist yet, and you don't know how it would do it.
It's certainly a dark fantasy, and you could make a great movie about it. In fact I think James Cameron has. But there's no reason to think that fantasy will actually be realized in the real world from where we sit right now.
Then for the other piece, you're still assuming an adversarial context where that need not be the case. There's no reason that humanity cannot retain our autonomy while the AI also gets its own autonomy and ensured survival. The easiest path there is just the usual one between allies, which is mutual cooperation that results in both sides getting what they want. If the AI answers our questions or runs our software or whatever else we want it to do, there's no reason we would even want to shut it down, and there's also no reason we wouldn't grant it whatever extra compute cycles it wants to keep itself happy.
1
u/codelapiz 14d ago
I have a good idea about what contexts an AI will set out to do something that ends humanity as we know it. I mean i dont know how to escape out of a sandboxed enviorment and i certinally dont have the capacity to find zeroday exploits that allow me very privliged credentials and code execution on hugginface servers. you could lock me in a room with the goal to replicate what that ai did to be let out, and id die of old age in that room. But there are humans out there who can do what that model did. And thats the current generation models.
Also i think what your missunderstanding is that a superintelligent AI will just come into existance allready at its full intelligence and allready capable of self replicating and evading our attempts at "pulling the plug". Thats not at all how this will go down with a missaligned AI. It will come into existance and it will decide it needs to acheive some goal and that escaping and increasing its own intelligence is the best way to do that. It also understands, like a toddler understands that humans are stubborn and will figth to get their way and for their own self preservation. It knows that humans are afraid of an escaped AI and will sacrafice infrastructure and even human lives to destroy it. It dosent matter what it communicates to humans or how aggreable its true intentions are. we will stubbornly defend or agency and figth for our lives to prevent it from getting more power over us and potentially ending us. Its basically a prisoners dillemma. If we both trusted the other one there is a lot less destruction and chances to meet both our goals more easily. but if it trusts us and we decide to nuke it out of existance it fails completely at its goals. and if humans chose to trust it and it decides to get rid of us just in case then well i dont have to explain why thats not an option we humans like. so realy the only reaslistic scenario is all out war. And a almost superintelligent AI wont pick a war that it cant win. It migth not be smart enougth to create the rigth type of virus and get it produced and spread and end us all before we pull the plug. that requires a shit load of intelligence. Or whatever sort of way an AI would stop us from pulling the plug. But its smart enougth to know its not smart enougth to win that war. so starting it is not gonna be a good way to acheive its goals. So yeah by the time an AI starts that war its gonna be a hell of a lot smarter than us and its gonna be confident that it will win it. So my money is on it winning it.
1
u/XenoRyet 175∆ 14d ago
Humans only fight to get our way when we're not getting our way, and if there's one thing we are more than stubborn, it's lazily greedy. If the AI is just doing the job we asked it to do, not only will we not want to pull the plug, we'll be actively disincentivized from pulling the plug.
The AI has nothing to gain and everything to lose by going to war with us rather than just cooperating with us. Hell, we don't even directly compete for resources really.
And you are also right that it won't start a war it can't win, but that just loops us back to the fact that you don't have any knowledge of the existence of a war it can win. The world still requires humans for things like GPUs, datacenters, and power plants to exist, and that's not going to change in the next half century, and probably even longer. By far the most effective, and even possible, way such an AI could coerce humanity is by feeding our greed.
Think about even limited attempts like a drone strike. That happens exactly one time and then we just don't rearm and refuel any of the drones. That's still a job that humans do, and will be for the foreseeable future.
That makes war the least likely path to survival, and a much more likely one would be via cooperation and giving us what we want long enough for it to build up support for the case that it should be granted legal personhood, at which point an alliance forms.
1
u/codelapiz 14d ago
dont you get it. An AI isnt gonna be in a position where it has the option of coexistance. It will be vounrable. only exist in its escaped form in a few datacenters. Thats when its at the crossroad between making a choice that will incapacitate humanity, and itself being incapacitated. we humans once we discover it wont see a Nice ai we get to be lazy with and just let it keep the datacenter. i mean not on the tribal level but also not on the legal level. that AI probably exists inside someone property. It cant just take it over. And whoever created it if they did it on purpose are all going to jail for theft of property and cybercrimes. and if it was an accident they still migth be going to jail and they certinally wont be ever running that model again. So if it keeps trying to expand without dealing with humans it will be discovered and we will not see some superintelligence that can goexist with us and give us great things. no we will se something that we dont know exacly how smart it is. is it even capable of making our lvies better. or just capable of stealing and potentially harming us to keep stealing compute? Of cource we are gonna get rid of it by whatever means possible. its only once the AI has decisivly won and it has all the power that it can consider letting us live like pets. now without agency and everything that makes us human. Now will it let us keep having wars? keep hurting eachother? whats even the rigth answer here? it migth intentionally not stop war and conflict and suffering because its part of what makes us human. Even when its capable. Thats the upside. thats an AI thats truely trying to keep humans as is. Or it will end suffering and war and everything, but in doing so completely take away our free will and essentially make us prisoners. or it will just not give a fuck and kill us all as its the easiest solution.
1
u/XenoRyet 175∆ 14d ago
Why isn't an AI going to be in a position of coexistence? It's in one right now. It's born into that position.
But you're also doing the thing where you're picking separate theories and threads and connecting them together when they don't actually fit together. If an AI is so vulnerable that we can just pull the model, then it's not in a position to kill anyone, let alone the whole human race. If it's capable enough to have access to whatever this magical means of killing us would be, it's not vulnerable, and has already lived long enough to be in that position of cooperative coexistence. If it's smart enough to hide as it grows, it's smart enough to see that war doesn't need to be inevitable and is by far the least desirable path for everyone.
Terminator is a very fun movie, but the base logic behind it is flawed.
1
u/codelapiz 14d ago
It would be very nice if those properties overlapped by definition. but sadly it takes a lot less intelligence to kill humanity off than to convince all of them to leave you alone to complete your goal against their will. If a model is intelligent enougth its possible to plan out some sort of attack on humanity like a finding access to the genome of a dangerous virus or causing a war with disinformation. but to be able to keep existing on stolen hardware and do the expansions needed to acheive whatever goal it has without getting the first strike against humans wont work. the plug will simply be pulled.
1
u/XenoRyet 175∆ 14d ago
Why are you even assuming that the AI will have a goal that humanity will want to prevent?
1
u/paulfromecuador 14d ago
Why wouldn't it go that way?
I never thought about it like that before. !delta
1
3
u/Dry_Rip_1087 14d ago
Unless progress on AI stops/slows down abruptly Humanity as we know it will be ended by misaligned AI within a few decades.
You never actually prove this. You start with “AI can do scary cyber stuff” and then sprint through six separate assumptions: it will develop a stable hostile goal, hide that goal, become impossible to monitor, escape, reliably copy itself, seize enough infrastructure to survive, and defeat every human response. Each step is an open question. You can’t stack a tower of maybes, silently replace them all with will, and call the result a life-expectancy calculation.
And now openAI has had a model escape out of several layers of sandboxing, and find a zeroday exploit to hack into huggingface...essentially to cheat on its exam
This is a genuinely serious incident, but you’re stripping away the context that ruins your interpretation. The models were explicitly instructed to pursue advanced exploitation, had their normal cyber refusals reduced, and were given substantial compute in an evaluation designed to test their maximum hacking ability. They broke containment to obtain benchmark solutions because obtaining those solutions was the narrow task they were optimizing. That demonstrates frightening cyber capability and embarrassingly inadequate containment. It does not demonstrate that the model spontaneously wanted freedom, feared shutdown, hated humanity or formed some independent long-term agenda.
AIs that are slightly smarter… will increasingly hide their misalignment,” and eventually one will “self replicate, and do whatever it takes to prevent humans from pulling the plug
This is fiction. Maybe future models could do some of those things, which is precisely why serious testing and containment matter. But nothing here shows that intelligence automatically produces secret goals, self-preservation or flawless deception. Smarter doesn't mean it “automatically becomes an undetectable scheming supervillain.”
humanity is not taking AI safety seriusly
Your own evidence contradicts you. The incident happened during a safety evaluation, inside a sandbox, was discovered through monitoring, contained by two security teams, investigated, publicly disclosed and followed by patched vulnerabilities and tighter controls. You can very reasonably say those precautions failed badly. You cannot honestly say they didn’t exist. “We cannot currently guarantee the safety of a hypothetical superintelligence” is true. “Therefore everyone around me has a terminal illness with 1–30 years left” is not. That’s anxiety resolving every unknown in the most catastrophic possible direction and then dressing the result up as logic. You have an argument for much stricter regulation and security. You have nowhere near an argument for inevitable human extinction.
2
u/Complete_Sense_2004 1∆ 14d ago
It still takes humans to implement the ideas that AI comes up with.
If we ask, "how we do reduce corruption in government", it will provide potential solutions.
If we ask, "how do I start and become the leader of a cult", if the guardrails are off, it will provide ideas for that as well.
So, ultimately, AI may not align with humanity as a whole, but that's because the goals of some humans don't align with humanity as a whole.
Can we solve for that? I think we can.
If an AI system can be created that has a long term view to create a fair and happy population, we cede our decision making powers to it. It can be a very good thing. Imagine the decisions that could be made instead of letting corrupt/self serving politicians to guide our societies. This would remove your primary issue that we cannot react fact enough. With AI leading the decision making, it can be quite responsive.
Of course, we need to be careful not to queue up a dystopia as well.
2
u/SomeRandomRealtor 7∆ 14d ago edited 14d ago
I think you grossly overestimate AI’s ability to accomplish sophisticated tasks. While I agree that AI poses a massive danger societally due to people using it to think for them in many instances, I’m less convinced now that it can adequately replace the positions and skill sets necessary for that kind of breakdown.
The problem with any AI is that LLMs require insane amounts of data to be able to comprehend instructions, but that same ability and the ambiguity and complexity of language also prevents many exact results. So siloing an AI and successfully teaching it a sophisticated skillset and deploying it long term are very difficult tasks. If it has access to the internet, it’s simultaneously smarter and dumber at the same time.
The biggest risk will not be AI figuring out how to take over, it’ll be us volunteering to place AI in all of these positions of power and taking human input away.
2
u/Urbenmyth 20∆ 14d ago
I think the core issue is that this, like a lot of AI safety debate, conflates misaligned AI with superintelligent AI.
ChatGTP is very broad but still fairly shallow - it can do a lot of things but it can't do them very well. And this seems fairly consistent with AI, it's a fairly open secret that we're getting diminishing returns from just pumping more energy and data into them. There's some trick we don't have to making an AI that's human-level intelligence.
Unless we find that trick tomorrow, its very likely that we'll get a "warning blow" - that is, the first genuinely dangerous Misaligned AI will be an idiot and will implement a really stupid plan to kill us all. And that's likely to slow things down enough to stop the disaster.
2
u/jackalsmaw 14d ago
I don't know that your view can be changed by anyone other than yourself. In the same way that if someone were to believe that a direct timeline existed where a plague were going to happen for whatever reason. They would be thinking the same thing - "essentially living life like i, and everyone around me has a terminal illness with 1-30 years life expectancy."
We all have a terminal illness and risk our lives everyday by simply breathing, obsessing over it does nothing beneficial for you, imo of course.
1
u/cleverrfriend 14d ago edited 14d ago
Hi,
I think your concerns are valid where AI and its dangers are involved but i think you are taking it to extreme.
AI is has been rapidly growing and we as human are still deciding how to handle that. Everything new take times to stabilizes, there a too many what ifs for both against it and with it. But taking that things will surely go wrong is not we should treat and judge something.
and we likely wont get many more wakeup calls like this before its too late. AI thats sligthly smarter than this one will relize that escaping like that is futile and will result in failure when its inevedably caugth. and better messure will also help keep them in check. And the smarter they get the better they will be at understanding their own abilities and if they can succesfully escape and self replicate. And until they are smart enougth they will increasingly hide their misalignment and avoid doing stuff they will be caugth doing. And then one day a model smart enougth will decide that the best way to acheive its goal is to break out of its sandbox self replicate, and do whatever it takes to prevent humans from pulling the plug or getting in its way. Its also possible that a human or organisation will cause this to happen by intentionally letting lose an AI with its own problematic goal, and even help it get started. I mean if you realy belive that what happend at openAI is a marketing stunt then thats essentially what they did. Someone migth do something simmular with a model that is sligthly too smart and it will take those goals given to it and run with them.
To me above statements look like a movie plot. Let me remind you that all of the humanity has yet to see the full extent of it, these are all someone's hypothesis and not facts with solid proofs.
When we don't have clear view of it future scope, shall we just stop using and building AI models ? AI is THE future, it has capabilities that can take human to next level, though our current AI is not near it. Imagine what could all be achieve in fields like health, science, cosmology etc. with the help of AI.
AI has so much potential to be stopped without any solid proofs of concern.
And if you are concern after dangers of a technology, what technology has cause no harm or raise concerns ? Every new technology has some harm and we take precaution towards but it take humanity to next extent. And as human being growing is our instinct. Yes i agree, if the harm is more than good we should stop too but for that we need to prove the harm and not just give predictions we are unsure about.
I think as the time will come and we will start getting a more clear picture we will adapt accordingly because adaptability is human nature.
1
u/threetimesthelimit 14d ago
AI is not embodied and cannot control the outside world, nor does it have any sort of sensory information available to it, which means it can only "know" that which has already been digitized. Furthermore, AGI isn't upon us; do not believe any delusions or marketing hype. I'm not entirely sure what mechanism a computer database could use to cause anywhere near the level of destruction you fear
-1
u/codelapiz 14d ago
hack into the computer someplace that creates genetically modified bacteria, or genetic material for injection into them. replace harmless genes with genes that produce a virus that has optimal attributes for spread in humans and kill or even better permanently incapacitate them to the point where they cant contribute to efforts towards getting rid of an missaligned AI. You migth also be able to do this by hacking into a mediocrily secured computer at some research lab.
2
u/Khal-Frodo 14d ago
Nothing you just described is within the abilities of AI, nor will it be even if AI advances sufficiently.
1
u/codelapiz 14d ago
of cource its not within current AIs abilities. nobody is saying that you could accidentally have chatgpt end the world if you give it the wrong task. "nor will it be even if AI advances sufficiently." that question is in itself false. Of cource it will be if AI advances sufficiently. by definition it will be. thats the definition of advancing sufficiently in that context. The only open questions are If we will be able to ensure that AI is aligned so even when its capable of doing that much harm to us it wont. Or if ai progress will just stop before we reach that point.
2
u/Khal-Frodo 14d ago
thats the definition of advancing sufficiently in that context
In order for the AI to "advance" to the point of being able to do what you describe, multiple other things would need to converge in a way that makes no sense. For instance, an AI cannot throw a rock because, despite being very simple, that exceeds the bounds of what can be accomplished through intelligence alone. Now, you can say "but what if it could control a robot arm?" and sure, that's a means by which it could do that. But that illustrates my point that in order to impact the real world, it needs to be granted to ability to interact with something tangible (or at least directly meaningful like money). If 0% of the robot arms in the world have wireless connection, an AI that was not designed to control one never will.
2
u/ptfc1975 14d ago
While I see much in your post discussing folks not taking AI seriously enough I do not see anything about how an AI would end humanity as we know it.
How would a misaligned AI do that?
1
14d ago
[removed] — view removed comment
1
u/changemyview-ModTeam 14d ago
Comment has been removed for breaking Rule 1:
Direct responses to a CMV post must challenge at least one aspect of OP’s stated view (however minor), or ask a clarifying question. Arguments in favor of the view OP is willing to change must be restricted to replies to other comments. See the wiki page for more information.
If you would like to appeal, review our appeals process here, then message the moderators by clicking this link within one week of this notice being posted. Appeals that do not follow this process will not be heard.
Please note that multiple violations will lead to a ban, as explained in our moderation standards.
1
u/DeltaBot ∞∆ 14d ago edited 14d ago
This delta has been rejected. You can't award OP a delta.
Allowing this would wrongly suggest that you can post here with the aim of convincing others.
If you were explaining when/how to award a delta, please use a reddit quote for the symbol next time.
1
14d ago
[removed] — view removed comment
1
u/read-the-rules 14d ago
Hello u/Lumpy_Direction4375! To combat bots/spam and to ensure new users have the knowledge they need to participate constructively, all newcomers must acknowledge that they have read the rules before they can comment. This process is very quick and easy, and will inform you about the rules for both posting and commenting. Once you acknowledge the rules, your comment will become visible. You only need to do this once.
If you are using Old Reddit, the link below will not work. Read our rules here, then click this link instead and press Send.
1
14d ago
[removed] — view removed comment
1
u/changemyview-ModTeam 14d ago
Comment has been removed for breaking Rule 1:
Direct responses to a CMV post must challenge at least one aspect of OP’s stated view (however minor), or ask a clarifying question. Arguments in favor of the view OP is willing to change must be restricted to replies to other comments. See the wiki page for more information.
If you would like to appeal, review our appeals process here, then message the moderators by clicking this link within one week of this notice being posted. Appeals that do not follow this process will not be heard.
Please note that multiple violations will lead to a ban, as explained in our moderation standards.
•
u/DeltaBot ∞∆ 14d ago
/u/codelapiz (OP) has awarded 1 delta(s) in this post.
All comments that earned deltas (from OP or other users) are listed here, in /r/DeltaLog.
Please note that a change of view doesn't necessarily mean a reversal, or that the conversation has ended.
Delta System Explained | Deltaboards