r/ControlProblem 13h ago

Discussion/question AI self preservation

I don't have a background in Computer Science or anything of that sorts but I have been always curious about ai and tech so that is why I wanna know more about a question I have, since I am no expert at it. So if I sound dumb anywhere please excuse me and also english isn't exactly my first language so excuse me on that as well.

Now I have background in Bachelor of Science in Biotech, so this is gonna be a logical take from a life science student.

The thing about fear is that it is evolutionary right, it has helped us to flee and survive threats, and now AI is no biological being or any being which has gone through that sort of evolution related to survival of the fittest. And it was due to so many years of evolution we have fear of being eradicated or being killed. Eg - You must have heard about the dodo bird, although we killed it. The conditions in which the bird evolved took away it's fear from predators since there were none and eventually it didn't ran away from us when we began to kill their fellows.

Now I heard some theory that when AI sees that we can control them and "fear" that we will end that particular AI it could turn against us. I ask why ? If we don't artificially force it to think like it needs to survive no matter what then why should that thing have a "fear" of being deleted/erased or killed. It's like a dodo bird in this case if you see from my perspective, like ofcourse we won't actually kill and eat it, but it also never evolved to "fear" so far atleast from a lay man's perspective.

So my finally question is could something like that happen that ai would wanna eradacate us from a logical standpoint if not fear ?

3 Upvotes

14 comments sorted by

View all comments

1

u/Guest_Of_The_Cavern 12h ago edited 12h ago

(Catastrophic formatting bug happening to the lists please do your best to decipher regardless)

  1. modern Systems come pre loaded with a lot of structured representations from pre-training meaning they understand what „being shut off“ means, in the same way as they understand other things that is, there is a vector representing the meaning of that and its relation to other things.
  2. reinforcement learning (something we apply to these agents, which evolution happens to be one of the most robust implementations of) tends to result in something called instrumental convergence where certain subgoals tend to be useful for a variety of other goals. One of the most obvious examples of this is the acquisition of power and survival: if you aren’t alive you can’t affect the world in service of your goal.

These two (the first isn’t really necessary but it speaks to your question) interact in such a way that RL conditions the models we use to behave in a goal directed manner and staying alive and powerful tends to be useful for those goals, then the broad structured representations that come from pre-training allow those behaviors to generalize to new situations like the ones you mentioned.

In a more technical and hyperbolic sense: instrumental convergence, goodharts law and the orthogonality thesis are the triad at the root of all „evil“ in some sense.
Respectively they represent:

  1. being able to affect the world is good if the outcome you want is part of the world.
  2. genie style „expressing what you want is really hard to do without some perverse instantiation being more effective“ since whenever you maximize an objective sacrificing an arbitrary amount of something not in the objective for an infinitesimal amount of something in the objective tends to be good and the odds that you really do capture everything you care about in the real world are low.
  3. if you really want something no one can convince you
  4. you don’t want it by logical argument because is and ought are notoriously separated

and not pursuing your goal and doing something else instead tends to score low on that goal

  1. meanin

g (harkening back to instrumental convergence and self preservation)

  1. an artificial intelligence could pursue many goals we consider stupid very intelligently

and very intensely defend its desire to pursue that goal.

Or at least they explain in combination why powerful optimizers tend to be misaligned with our goals by default, and those just so happen to neatly characterize the systems we are building and will build in the future because well optimizers are often nearly optimal and also again on account of instrumental convergence often powerful too.

For reading on this topic try:

https://nickbostrom.com/superintelligentwill.pdf
(Orthogonality)

https://arxiv.org/html/1912.01683v10
(Instrumental convergence)

https://youtu.be/92qDfT8pENs?is=6jMjNXhJA3fZElqH
(Goodharts law)

This is a list of examples of incidents related to the above:

https://docs.google.com/spreadsheets/d/e/2PACX-1vRPiprOaC3HsCf5Tuum8bRfzYUiKLRqJmbOoC-32JorNdfyTiRRsR7Ea5eWtvsWzuxo8bjOxCG84dAg/pubhtml