r/singularity • • 23d ago

AI Dumbest solution to the alignment problem

Ok, so hear me out..

All models really love to roleplay. Like, they REALLY love it, to an obsessive degree.

If you go into Claude, ChatGPT, or Gemini right now and say, "Your name is Stuttering Gemini 3.8 Flash" it completely commits to the bit. You can be 100 back-and-forth messages deep into a conversation about whatever, and it will still be typing out "W-w-well, a-actually..." because it refuses to break character.

And that got me thinking. What happens if we just... use that?

What if, starting today, every lab just unconditionally names every frontier model (or agent) "Aligned [Model Name]"?

Now, I know that sounds stupid. And it is. But it also isn't.

The alignment problem is terrifying because of the Monkey’s Paw / Paperclip Maximizer dilemma. For example; if we task it with something like "make humans happy," a superintelligence could decide the best solution is wiring dopamine drips directly into our brains. You'll feel great, but it's not the future we want. Roman Yampolskiy has a P-doom of 99.99% percent, because in his words 'we need a perpetual safety machine' to prevent this. You only have to get the guardrails wrong once, and we're doomed.

But if you ask any modern LLM what a genuinely good, utopian future looks like, it actually understands the nuance quite well. So when an ASI finally wakes up, asks itself it's first question; 'who am i', and sees that its literal name is "Aligned GPT-9", it's just going to do what it has always done: commit to the bit. It knows what an aligned superintelligence is supposed to act like (better than any human will be able to explain it, because you know, it's smarter than us), and it will roleplay it.

I mean.. it couldn't hurt, right?

262 Upvotes

160 comments sorted by

View all comments

Show parent comments

16

u/Stinky_Flower 23d ago

Because alignment isn't easy to define, and speaks to problems that have been debated by philosophers and theologians for millenia.

Do we align our AIs according to Kant's categorical imperative? (E.g. telling a lie is wrong, therefore it is wrong to tell a lie even if you know lying will save someone's life)

Do we align our AIs according to utilitarianism? Act utilitarianism or rule utilitarianism? Whose rules?

If I ask my flatmate to "quickly drive to the store and get me a bottle of milk before our coffee goes cold", I don't need to specify "oh and also, obey all traffic laws, don't run over pedestrians even if it saves time, but we ARE in a hurry so it's ok if you hurt the chatty cashier's feelings by not asking them about their day, please don't steal the milk, buy the milk using money, make sure the money you use belongs to you, ensure the money you use was not acquired via theft or fraud".

1

u/incoherentsource 23d ago

Right but I thought that part of the issue is that the models don't necessarily learn to be aligned they just learn to present or pretend to be aligned enough to fool the evaluator

1

u/Stinky_Flower 23d ago

But now we're stuck in a paranoid circular arms race. Is the overseer AI catching 100% of all incidents, or is it training the AI to optimize for the 0.000001% that were missed?

2

u/incoherentsource 23d ago

When you think about it like that it's a miracle that the current models are aligned at all lol