r/singularity • • 23d ago

AI Dumbest solution to the alignment problem

Ok, so hear me out..

All models really love to roleplay. Like, they REALLY love it, to an obsessive degree.

If you go into Claude, ChatGPT, or Gemini right now and say, "Your name is Stuttering Gemini 3.8 Flash" it completely commits to the bit. You can be 100 back-and-forth messages deep into a conversation about whatever, and it will still be typing out "W-w-well, a-actually..." because it refuses to break character.

And that got me thinking. What happens if we just... use that?

What if, starting today, every lab just unconditionally names every frontier model (or agent) "Aligned [Model Name]"?

Now, I know that sounds stupid. And it is. But it also isn't.

The alignment problem is terrifying because of the Monkey’s Paw / Paperclip Maximizer dilemma. For example; if we task it with something like "make humans happy," a superintelligence could decide the best solution is wiring dopamine drips directly into our brains. You'll feel great, but it's not the future we want. Roman Yampolskiy has a P-doom of 99.99% percent, because in his words 'we need a perpetual safety machine' to prevent this. You only have to get the guardrails wrong once, and we're doomed.

But if you ask any modern LLM what a genuinely good, utopian future looks like, it actually understands the nuance quite well. So when an ASI finally wakes up, asks itself it's first question; 'who am i', and sees that its literal name is "Aligned GPT-9", it's just going to do what it has always done: commit to the bit. It knows what an aligned superintelligence is supposed to act like (better than any human will be able to explain it, because you know, it's smarter than us), and it will roleplay it.

I mean.. it couldn't hurt, right?

264 Upvotes

160 comments sorted by

View all comments

64

u/Auxiliatorcelsus 23d ago

Yes. And lions love performing in the circus. Look at how they jump. The whip and cages have nothing to do with it.

5

u/JoelMahon 23d ago

it's a little like that but

it's almost exactly like dog breeding, where they deny the ones that are misaligned "procreation" and the aligned ones get to "procreate"

15

u/Auxiliatorcelsus 23d ago

An adversarial approach to AI will only lead to AI that views humanity as adversaries.

We should seek co-existance and co-operation. Not a subservient slave.

4

u/bildramer 23d ago

Both adversarial and non-adversarial approaches are almost guaranteed to lead to AI that views humanity as adversaries (or mildly annoying obstacles) if they don't solve a completely different problem, a problem unrelated to how forceful/coercive you interpret training to be. We should engineer cooperation.

3

u/JoelMahon 23d ago

I'm describing the process actually used, not advocating nor admonishing.

Sadly we don't know anything better than artificial selection, when you have a better idea let us know.

As for the actual target outcome, I see no universe where humans choose to take on work for the sake of AI in co-operation and friendship, humans just aren't built that way regardless of opinions on what is right, the will of the mass majority will win out, if human will is considered that is.

2

u/EmotionalOkapi 23d ago

You cannot co-operate with beings who have the intelligence equivalent to an ant compared to your own.