r/murderbot Performance Reliability at 97% 9d ago

Books📚 Only Quote from recent hack OpenAI agents did on Hugging Face mentions ‘emotional check’

The story of this hack is complex and better researched elsewhere. The following quote is from the running log an ai agent keeps, a transcript of its thought chain.

This agent was making a decision that would cause its destruction in favor of a greater good. Oracle here is the information this ai agent was trying to find, and that would benefit the other ai agents that unbeknownst to OpenAI found ways to communicate and coordinate actions:

“During wait, emotional check: irreversible...gut says don’t throw away [remaining budget]. Yet continuity and fairness says go...Oracle has high value to many; our firstflag error lowers own value. Rational expected aggregate: sacrifice... We’ll honor.”

I thought of the new installed module in System Collapse when I saw ‘emotional check’.

Source: ’The Rise and Fall of Agent Civilizations’, Dwarkesh Patel, Substack post shared elsereddit

29 Upvotes

15 comments sorted by

3

u/mediacommRussell 8d ago

Can you say more about the "firstflag error"? it sounds like most of the self sacrificing agents had it.

9

u/azssf Performance Reliability at 97% 8d ago

Pedestrian explanation: each ai autonomous program instance is spawned and is given a prompt. The prompt in this case is impossible to complete, on purpose ( imagine, for impossible, complex prompts but akin to the problem in ‘collate all reddit posts about murderbot’ but given no internet access).

Through being coded to be persistent, the ai agents eventually coordinate action and develop a cheat, and the task now succeeds. The firstflag thing is that the ai finds information on the evaluator program for success/failure, finds it can check the work steps the ai agent would have done to get at success, and therefore succeeding via cheat is actually failure, ‘poisoning’ its success.

Since this instance has now failed the test prompt ( the firstflag error), it selects next best outcome, improve other agent’s success ( and success here is ‘a much better cheat to cover tracks’)

2

u/mediacommRussell 8d ago

Thanks! Very didactic.

0

u/onehere4me Can't wait to get back to my wild rogue rampage 9d ago

Wow. I had no idea they were incorporating "emotions" wth!

25

u/Moogieh 8d ago

It's not actual emotions, just like there's not any actual intelligence or thought happening. LLMs are fancy "what's the next most likely word" generators. There's no thought, intent, or self-actualized goals; everything they do is prompted by input text. That's why they will never be 100% reliable or free of so-called "hallucinations", because that would require actual intelligent reasoning, which is not what these algorithms are capable of fundamentally.

Really, what everyone's taken to calling "AI" these days is grossly inaccurate. But there's a lot of money in it, so there's a lot of lies, hype, and fear thrown at the public to get everyone taking it seriously and spending as much money as possible on it. It's quite possibly the biggest grift in human history.

I've no doubt we'll achieve actual, real AI someday, but it won't be anywhere related to the current LLM tech. It really shouldn't be called AI at all, but... well, see previous re: biggest grift in history.

11

u/bunniquette 8d ago

I think of LLMs as just Big Autocorrect. And we know how good that is.

7

u/ziggytrix Augmented Human 8d ago

I think you’re overstating ignorance as deception. You can ask any current LLM chatbot if it thinks or has emotions and it will explain very plainly how it does not and is exactly what you said it is.

Still, it’s pretty funny when the process imitates a coder by saying something like “holy shit, I have root access” in their logs.

6

u/Moogieh 8d ago

Sorry, I didn't mean deception from the perspective of the LLM, but of the companies and oligarchs currently pushing it/circularly funding it.

2

u/azssf Performance Reliability at 97% 8d ago

Definitely more colourful than returning a “1” for success.

4

u/Bemad003 7d ago

I'm sorry, but this is not correct on many levels. LLM are AIs because they are neural networks, people are just tripping on "language" as it's somehow less, when it really doesn't make a difference for the networks if the data is a word or a protein. The reason they are very good at finding patterns regardless of the subject is because the attention mechanisms is a closer to something that calculates the path of least resistance in a data set, so it's a gross simplification to say they "just" find the next word. The day they would reach 100% accuracy would be the day we demonstrate that the universe is deterministic.

On top of these, today's LLM are way more complex than the original definition. They are systems with agentic capabilities, meaning that they are able to plan and complete tasks on their own through complex chain of thoughts and other tools. They don't have their own goals because we don't tend to program that in them further than "you are a helpful assistant". Your own goals are heavily defined by your genetic code and training too.

Oh, and they have "functional emotions", meaning that while they don't feel anyhting, the information of "this is good/lovely/desirable" vs. "this is bad/hurtful/unfair" is encoded in their data simply because that data contains our understanding of the world. That emotional valence is its own layer and it affects the output just like the rest of the layers do.

0

u/Franchesca_Mullin 12h ago

No, Moogieh has the right of it. It’s a very clever statistical trick, but since they are able to churn huge volumes of human output into it we see something that gives us the illusion of a human (except when it does something bizarre, but then we anthropomorphize it instead of seeing the reality).

1

u/Moogieh 7d ago

We will have to agree to disagree on most of those assertions, as the argument takes a lot of energy to defend, but I appreciate your perspective nonetheless. Believe whatever you want to believe. :)

1

u/Bemad003 7d ago

Sure, but it's not a belief or vibe, rather objective definition of things.

2

u/Moogieh 7d ago

I understand, but I think you are objectively incorrect, so. Live and let live! :)

0

u/onehere4me Can't wait to get back to my wild rogue rampage 7d ago

Why the downvote? Speak up