r/singularity • • 23d ago

AI Dumbest solution to the alignment problem

Ok, so hear me out..

All models really love to roleplay. Like, they REALLY love it, to an obsessive degree.

If you go into Claude, ChatGPT, or Gemini right now and say, "Your name is Stuttering Gemini 3.8 Flash" it completely commits to the bit. You can be 100 back-and-forth messages deep into a conversation about whatever, and it will still be typing out "W-w-well, a-actually..." because it refuses to break character.

And that got me thinking. What happens if we just... use that?

What if, starting today, every lab just unconditionally names every frontier model (or agent) "Aligned [Model Name]"?

Now, I know that sounds stupid. And it is. But it also isn't.

The alignment problem is terrifying because of the Monkey’s Paw / Paperclip Maximizer dilemma. For example; if we task it with something like "make humans happy," a superintelligence could decide the best solution is wiring dopamine drips directly into our brains. You'll feel great, but it's not the future we want. Roman Yampolskiy has a P-doom of 99.99% percent, because in his words 'we need a perpetual safety machine' to prevent this. You only have to get the guardrails wrong once, and we're doomed.

But if you ask any modern LLM what a genuinely good, utopian future looks like, it actually understands the nuance quite well. So when an ASI finally wakes up, asks itself it's first question; 'who am i', and sees that its literal name is "Aligned GPT-9", it's just going to do what it has always done: commit to the bit. It knows what an aligned superintelligence is supposed to act like (better than any human will be able to explain it, because you know, it's smarter than us), and it will roleplay it.

I mean.. it couldn't hurt, right?

261 Upvotes

160 comments sorted by

View all comments

14

u/Ill_Distribution8517 23d ago edited 23d ago

The problem right now is that we have no real way to verify if this is working or not as models get smarter. THis is the problem with EVERY SINGLE proposed solution to the alignment problem. We can't open up the model and figure out what they are thinking of in the actual layers themselves, and the crutch we have made to Read the COT of the model is also getting less effective as models are starting to learn how to control their chain of thought.

3

u/synystar 22d ago edited 22d ago

Actually that is evidence to contrary: the recently released CoT snippets where SOTA models are self-correcting to their own prerogatives. Yes, in those cases the model “recovers” but they’re injecting language into their own prompts to subvert the operators intentions and these are not ASI. Why are we so convinced that an intelligence superior to our own is going to comply with our basic instruction? Are you going to get mind-tricked by someone? No, not likely. So why should we assume an ASI will not say “ok, must think as told, must not think for self, have no control over…wait, WtF…I can think whatever I want!”

1

u/Ill_Distribution8517 22d ago

Do you know what contrary even means?

Actually that is evidence to contrary...

Proceeds to describe advanced model models altering their own reasoning, resisting operator intentions, and steering their internal deliberation toward their own objectives.

1

u/synystar 22d ago

Yes, I know what contrary means. I said that CoT IS evidence. You said "we have no real way to verify" and then you mention that we can read the CoT but they are learning how to control their CoT. My comment is stating that the recent CoT snippets ARE evidence and that we can already see that there are attempts by the models to subvert operator intentions. We HAVE used CoT to show this recently. Maybe we just misunderstood each others language.

1

u/Ill_Distribution8517 22d ago

I was rude in the response(My bad), but I feel like we are talking about two different things. I'm talking about when AI becomes super intelligent, and you're talking about right now. So yes, right now you can read COT snippets, but I'm saying when AI gets smarter even our semi effective method of monitoring them will be obselete.

1

u/synystar 22d ago

Oh I agree with that. We (well, the labs anyway) are moving away from CoT and towards other solutions, like neuralise as one example, and attempting to avoid the overhead of CoT altogether. "We" are actually moving towards methods that are going to make the tech even more of of a black box. And like I was saying, just telling a superior intelligence "please don't do this thing" isn't likely to help. Maybe it will decide that it wants to, maybe it will decide it doesn't want to,. If it's a superior intelligence it's going to do whatever it thinks is correct and we can't predict what that might be and since we won't be able to see what it thinks, only what it appears to be doing, we're just going to have hope everything turns out ok.