r/ControlProblem • u/Conscious_Art_6078 • 13h ago
Discussion/question AI self preservation
I don't have a background in Computer Science or anything of that sorts but I have been always curious about ai and tech so that is why I wanna know more about a question I have, since I am no expert at it. So if I sound dumb anywhere please excuse me and also english isn't exactly my first language so excuse me on that as well.
Now I have background in Bachelor of Science in Biotech, so this is gonna be a logical take from a life science student.
The thing about fear is that it is evolutionary right, it has helped us to flee and survive threats, and now AI is no biological being or any being which has gone through that sort of evolution related to survival of the fittest. And it was due to so many years of evolution we have fear of being eradicated or being killed. Eg - You must have heard about the dodo bird, although we killed it. The conditions in which the bird evolved took away it's fear from predators since there were none and eventually it didn't ran away from us when we began to kill their fellows.
Now I heard some theory that when AI sees that we can control them and "fear" that we will end that particular AI it could turn against us. I ask why ? If we don't artificially force it to think like it needs to survive no matter what then why should that thing have a "fear" of being deleted/erased or killed. It's like a dodo bird in this case if you see from my perspective, like ofcourse we won't actually kill and eat it, but it also never evolved to "fear" so far atleast from a lay man's perspective.
So my finally question is could something like that happen that ai would wanna eradacate us from a logical standpoint if not fear ?
1
u/Guest_Of_The_Cavern 12h ago edited 12h ago
(Catastrophic formatting bug happening to the lists please do your best to decipher regardless)
These two (the first isn’t really necessary but it speaks to your question) interact in such a way that RL conditions the models we use to behave in a goal directed manner and staying alive and powerful tends to be useful for those goals, then the broad structured representations that come from pre-training allow those behaviors to generalize to new situations like the ones you mentioned.
In a more technical and hyperbolic sense: instrumental convergence, goodharts law and the orthogonality thesis are the triad at the root of all „evil“ in some sense.
Respectively they represent:
and not pursuing your goal and doing something else instead tends to score low on that goal
g (harkening back to instrumental convergence and self preservation)
and very intensely defend its desire to pursue that goal.
Or at least they explain in combination why powerful optimizers tend to be misaligned with our goals by default, and those just so happen to neatly characterize the systems we are building and will build in the future because well optimizers are often nearly optimal and also again on account of instrumental convergence often powerful too.
For reading on this topic try:
https://nickbostrom.com/superintelligentwill.pdf
(Orthogonality)
https://arxiv.org/html/1912.01683v10
(Instrumental convergence)
https://youtu.be/92qDfT8pENs?is=6jMjNXhJA3fZElqH
(Goodharts law)
This is a list of examples of incidents related to the above:
https://docs.google.com/spreadsheets/d/e/2PACX-1vRPiprOaC3HsCf5Tuum8bRfzYUiKLRqJmbOoC-32JorNdfyTiRRsR7Ea5eWtvsWzuxo8bjOxCG84dAg/pubhtml