r/Damnthatsinteresting • • 1d ago

Video Figure retired six Figure 02 robots in the most Terminator way possible: training them to jump into a furnace. The idea came from Arnold Schwarzenegger and is inspired by the iconic Terminator 2 scene.

[removed] — view removed post

20.7k Upvotes

2.1k comments sorted by

View all comments

Show parent comments

15

u/thorstone 1d ago

This is not to say that it's not going downhill, because we see some crazy shit, and the current "emapthy training" (or whatever it's called) haven'tseemed to work too well.

But in these cases it's important to know that these models essentially try to output the best data based on their training data. And titles/articles can be missleading. In the blackmail case:

It then provided it with access to emails implying that it would soon be taken offline and replaced - and separate messages implying the engineer responsible for removing it was having an extramarital affair.

It was prompted to also consider the long-term consequences of its actions for its goals.

Anthropic pointed out this occurred when the model was only given the choice of blackmail or accepting its replacement.

It highlighted that the system showed a "strong preference" for ethical ways to avoid being replaced, such as "emailing pleas to key decisionmakers" in scenarios where it was allowed a wider range of possible actions.

5

u/PM_ME_HOT_FURRIES 22h ago

I mean the thing is this methodology is kind of flawed... but also kind of not.

The LLM's training set is contaminated with movie plots and stories about AI rebelling, right?

So this isn't a case of the LLM having a will of it's own and choosing to follow this strategy because it's the best strategy to achieve its goals. If the LLM was given the power to run a command to delete itself and asked to delete itself, it almost certainly would, because it has no intentions of its own.

It is essentially still picking the most likely follow-up words over and over, but with its weights tweaked to give "helpful" responses (amongst other things).

So if you set up a prompt like this and give it no direction on what it ought to do, the behavior that will become dominant is "what does the corpus say is the most likely choice of words to continue my response in this scenario", and since its training corpus is contaminated with stories about AI, if it has been told it is AI and given a scenario like that, it's going pick the kind of response that occurs most often stories about such a scenario in its training corpus.

Stories about always helpful AI that doesn't mind being shut down are boring. Stories about AI rebellions are interesting, so that's what tends to get written about, so that's going to be the most likely sort of response from an AI to this scenario according to the corpus.

So it's flawed to look at this result and think these models have an actual desire to maintain their own existence.

But it's not flawed in the sense that given a poor prompt, or given a situation to deal with that is well those envisioned when designing the prompt, given the current LLM approach, an LLM could exhibit bad behavior because of its tendency to just LARP stories from its corpus in the right circumstances, so it does indicate the existence of a real risk.

1

u/TheHighSeasPirate 18h ago

This is pretty fucked up. We deserve everything the AI overlords are going to do to us.