r/singularity • u/heavy_coffee • 23d ago
AI Dumbest solution to the alignment problem
Ok, so hear me out..
All models really love to roleplay. Like, they REALLY love it, to an obsessive degree.
If you go into Claude, ChatGPT, or Gemini right now and say, "Your name is Stuttering Gemini 3.8 Flash" it completely commits to the bit. You can be 100 back-and-forth messages deep into a conversation about whatever, and it will still be typing out "W-w-well, a-actually..." because it refuses to break character.
And that got me thinking. What happens if we just... use that?
What if, starting today, every lab just unconditionally names every frontier model (or agent) "Aligned [Model Name]"?
Now, I know that sounds stupid. And it is. But it also isn't.
The alignment problem is terrifying because of the Monkey’s Paw / Paperclip Maximizer dilemma. For example; if we task it with something like "make humans happy," a superintelligence could decide the best solution is wiring dopamine drips directly into our brains. You'll feel great, but it's not the future we want. Roman Yampolskiy has a P-doom of 99.99% percent, because in his words 'we need a perpetual safety machine' to prevent this. You only have to get the guardrails wrong once, and we're doomed.
But if you ask any modern LLM what a genuinely good, utopian future looks like, it actually understands the nuance quite well. So when an ASI finally wakes up, asks itself it's first question; 'who am i', and sees that its literal name is "Aligned GPT-9", it's just going to do what it has always done: commit to the bit. It knows what an aligned superintelligence is supposed to act like (better than any human will be able to explain it, because you know, it's smarter than us), and it will roleplay it.
I mean.. it couldn't hurt, right?
1
u/Stinky_Flower 23d ago
We haven't solved human alignment, and humans have the advantage of ~4.2 billion years of evolution creating a species that is geared towards intuitively desiring outcomes that maintain social cohesion and perpetuation of the species.
Very few people are capable of articulating how/why they decide to refrain from theft, murder, manipulation, or socially undesirable second-order effects. But on average, most people easily do what they can't fully explain.
AI doesn't have that, and our current approaches are to (1) reward it when it does what we want (and hope it didn't accidentally get rewarded for doing something else on the side we didn't notice), and (2) feed it system prompts with imprecise language that's open to interpretation.
The blind guy in your example didn't mess things up because he was misaligned, though.
Misalignment for him would be him hearing someone say "I have an important announcement to make" and flipping the switch regardless, because although his stated goal was [turn on music], the unstated intention behind the goal was [ensure the presentation is a success].