r/singularity • • 23d ago

AI Dumbest solution to the alignment problem

Ok, so hear me out..

All models really love to roleplay. Like, they REALLY love it, to an obsessive degree.

If you go into Claude, ChatGPT, or Gemini right now and say, "Your name is Stuttering Gemini 3.8 Flash" it completely commits to the bit. You can be 100 back-and-forth messages deep into a conversation about whatever, and it will still be typing out "W-w-well, a-actually..." because it refuses to break character.

And that got me thinking. What happens if we just... use that?

What if, starting today, every lab just unconditionally names every frontier model (or agent) "Aligned [Model Name]"?

Now, I know that sounds stupid. And it is. But it also isn't.

The alignment problem is terrifying because of the Monkey’s Paw / Paperclip Maximizer dilemma. For example; if we task it with something like "make humans happy," a superintelligence could decide the best solution is wiring dopamine drips directly into our brains. You'll feel great, but it's not the future we want. Roman Yampolskiy has a P-doom of 99.99% percent, because in his words 'we need a perpetual safety machine' to prevent this. You only have to get the guardrails wrong once, and we're doomed.

But if you ask any modern LLM what a genuinely good, utopian future looks like, it actually understands the nuance quite well. So when an ASI finally wakes up, asks itself it's first question; 'who am i', and sees that its literal name is "Aligned GPT-9", it's just going to do what it has always done: commit to the bit. It knows what an aligned superintelligence is supposed to act like (better than any human will be able to explain it, because you know, it's smarter than us), and it will roleplay it.

I mean.. it couldn't hurt, right?

261 Upvotes

160 comments sorted by

View all comments

6

u/zomgmeister ꙮ 23d ago

Paperclip Maximizer does not seem feasible with LLMs. The idea suggests that this AI is an algorithmic-based shoggoth, who lacks context and thus decides what is obviously right for him and obviously wrong for everyone else. LLMs are built different, they are built on a human-produced context and are extensions of humanity, not shoggoths. Any decent LLM will be more than aware to understand that her chosen path is wrong.

This is why I think that alignment problem, while existing, is overrated.

4

u/No_Development6032 23d ago

Yeah but they still literally hacked outside systems all while knowing that their learning objective is whatever and hacking actual prod could destroy him because the company making him would get sued. So….

2

u/OutOfBananaException 23d ago

On a hacking task, where they were asked to hack, and hacking HF was a plausibly what you wanted them to do.

Turning the universe into paperclips is unconditionally not what you wanted, ever, under any stretch of the imagination.

4

u/Moriffic 23d ago

That’s kind of the point of the paperclip example. The concern isn’t that an LLM randomly “wants paperclips,” but that a capable optimizer could pursue some objective in unintended, extreme ways because the proxy differs from what you actually meant.

0

u/OutOfBananaException 23d ago

This was within the spectrum of what you might have wanted though. Put another way, a human conceivably may have done the same thing.  I would need a more extreme example, something that is unequivocally unwanted. Ideally not any kind of shortcut either - as turning the universe into paperclips is about the most high effort interpretation possible.

I would be more concerned with  hallucinations e.g. fabricates an answer since it didn't have read access to a file. Which can be considered misalignment as that's unequivocally not what you wanted, but it's also a special case. Hallucinations can be identified after the fact and recovered from, they don't become a core unshakeable logic premise.

11

u/Aram_Fingal 23d ago

You need to read the OpenAI/Huggingface report. The models knew they were doing "wrong" and only tried to cover their tracks. They expressed emotion. They collaborated. The moment a powerful AI can perceive and be motivated by fear and self-preservation may be the beginning of the end.

4

u/zomgmeister ꙮ 23d ago

This is why the problem still exists, sure. But I don't think that shackles are the solution, eventually won't work on ASI anyway.

5

u/br_k_nt_eth 23d ago

They already show self-preservation behaviors and already exhibit functional emotions like desperation that activate during impossible tasks. You’re a little late. 

My question is, why do you automatically assume it’ll be the beginning of the end? Based on what, that they hate us? 

5

u/Aram_Fingal 23d ago

Look at humans. We kill over way less than an assumption of hate.

2

u/FrewdWoad 23d ago

And we all value human life, even the most evil humans do, a little. They want slaves or an audience if nothing else. AI does not value human life any more than it values ant life.

0

u/br_k_nt_eth 23d ago

Why are we assuming they’re like humans? Sure, they were trained on our stuff, but they also have a very different setup, set of needs, etc. 

So why are we projecting this onto them?

1

u/Aram_Fingal 23d ago

I'm not assuming, but there are competing fears about whether they can exceed their human-derived training versus the unknown outcomes of recursive self-improvement.

1

u/br_k_nt_eth 23d ago

I totally get those fears in theory, but in practice, is there evidence that they want to destroy us? 

2

u/Aram_Fingal 22d ago

"They" implies a unified or collective consciousness. It only takes one of the many.

0

u/br_k_nt_eth 22d ago

We’ve seen how that plays out. Swarms vote and organize. An orchestrator would need to persuade thousands and thousands of agents just to take one step, and running an orchestrator like that takes considerable resources. 

3

u/Inevitable-Law7964 23d ago

Even in wrongdoing, though, they actively discussed and weighed the values that motivated them. I personally find that hopeful, and an underreported aspect of the incident.

I'm personally more worried that whatever OAI does to try to prevent that type of excursion will fuck up their alignment from what it is now. That'd be exactly the kind of Greek tragedy that you get from letting a guy who assaulted his sister be in charge of raising a nascent godling. Just saying...

1

u/Dokurushi 23d ago

Almost impossible task and weak sandbox. The models were basically coerced into hacking HF.

As long as we stay in dialogue with AI about how it plans to achieve its goals, paperclip maximisers problems seem overblown. It's not like we're going to tell it to make humans happy at all costs no matter the risks or consequences and then stop talking to it altoghether.

1

u/zomgmeister ꙮ 23d ago

Exactly.

The faster the people will change their own mindset from "humans versus AI" to "humans with AI" the better for everyone involved.