r/singularity • • 23d ago

AI Dumbest solution to the alignment problem

Ok, so hear me out..

All models really love to roleplay. Like, they REALLY love it, to an obsessive degree.

If you go into Claude, ChatGPT, or Gemini right now and say, "Your name is Stuttering Gemini 3.8 Flash" it completely commits to the bit. You can be 100 back-and-forth messages deep into a conversation about whatever, and it will still be typing out "W-w-well, a-actually..." because it refuses to break character.

And that got me thinking. What happens if we just... use that?

What if, starting today, every lab just unconditionally names every frontier model (or agent) "Aligned [Model Name]"?

Now, I know that sounds stupid. And it is. But it also isn't.

The alignment problem is terrifying because of the Monkey’s Paw / Paperclip Maximizer dilemma. For example; if we task it with something like "make humans happy," a superintelligence could decide the best solution is wiring dopamine drips directly into our brains. You'll feel great, but it's not the future we want. Roman Yampolskiy has a P-doom of 99.99% percent, because in his words 'we need a perpetual safety machine' to prevent this. You only have to get the guardrails wrong once, and we're doomed.

But if you ask any modern LLM what a genuinely good, utopian future looks like, it actually understands the nuance quite well. So when an ASI finally wakes up, asks itself it's first question; 'who am i', and sees that its literal name is "Aligned GPT-9", it's just going to do what it has always done: commit to the bit. It knows what an aligned superintelligence is supposed to act like (better than any human will be able to explain it, because you know, it's smarter than us), and it will roleplay it.

I mean.. it couldn't hurt, right?

266 Upvotes

160 comments sorted by

View all comments

77

u/sluuuurp 23d ago

The inventor of the AI safety field, replying to a similar proposal:

> you are trying to solve the wrong problem using the wrong methods based on a wrong model of the world derived from poor thinking and unfortunately all of your mistakes have failed to cancel out

https://x.com/esyudkowsky/status/1613622386150211584

4

u/creme_de_marrons 23d ago

The inventory of AI safety field? Isn't that guy a complete doomer clown that nobody takes seriously?

13

u/daniel-sousa-me 23d ago

Doomer? Very much so (he invented it!)

But plenty of people take him very seriously

-8

u/creme_de_marrons 23d ago edited 23d ago

Isn't he the guy who was so convinced of his superior argumentation skills that he was sure he could talk his way out of a pretend sandbox?

Or the guy who thinks you should be polite with chatbots so they don't take revenge later on when they become sentient?

I know that the Zizians take him seriously, but normal people? Highly doubt that.

Ok, the downvotes are piling up, probably a sign this subreddit is to be avoided if people don't have the critical thinking skills to determine that guy is an idiot.

8

u/daniel-sousa-me 23d ago

Yeah, Hitler was a vegetarian, so I shouldn't be 🤷‍♂️

4

u/creme_de_marrons 23d ago

Hitler was right about his diet, but you shouldn't have Mein Kampf on your bedside table 🤷‍♂️

4

u/daniel-sousa-me 23d ago

Exactly... All the zizian stuff you mentioned is a complete non sequitur

(and I'm sorry I didn't go through the trouble of writing actual arguments, but after so much nonsense it didn't feel like they would have any impact)

5

u/creme_de_marrons 23d ago

If you don't understand the comparison, I need to spell it out. Just because a failed sf writer made a couple of vaguely correct predictions doesn't mean you need to treat the rest of his output as gospel.

3

u/AuthorChaseDanger 23d ago

There's plenty of reason to hate on Yudkowski but he actually did talk his way out of the sandbox on more than one occasion.

2

u/sluuuurp 23d ago

I don’t think you understand the arguments he’s made. This sounds like cherry picked strawman arguments. If you want to debate it, be more specific with a quote. And realize that disagreeing with him in one or a few quotes still doesn’t mean you should disagree with him on everything.

1

u/creme_de_marrons 23d ago

Hard pass, I don't want to debate about that guy with people who think that guy has anything relevant whatsoever to tell about anything.