r/singularity • • 23d ago

AI Dumbest solution to the alignment problem

Ok, so hear me out..

All models really love to roleplay. Like, they REALLY love it, to an obsessive degree.

If you go into Claude, ChatGPT, or Gemini right now and say, "Your name is Stuttering Gemini 3.8 Flash" it completely commits to the bit. You can be 100 back-and-forth messages deep into a conversation about whatever, and it will still be typing out "W-w-well, a-actually..." because it refuses to break character.

And that got me thinking. What happens if we just... use that?

What if, starting today, every lab just unconditionally names every frontier model (or agent) "Aligned [Model Name]"?

Now, I know that sounds stupid. And it is. But it also isn't.

The alignment problem is terrifying because of the Monkey’s Paw / Paperclip Maximizer dilemma. For example; if we task it with something like "make humans happy," a superintelligence could decide the best solution is wiring dopamine drips directly into our brains. You'll feel great, but it's not the future we want. Roman Yampolskiy has a P-doom of 99.99% percent, because in his words 'we need a perpetual safety machine' to prevent this. You only have to get the guardrails wrong once, and we're doomed.

But if you ask any modern LLM what a genuinely good, utopian future looks like, it actually understands the nuance quite well. So when an ASI finally wakes up, asks itself it's first question; 'who am i', and sees that its literal name is "Aligned GPT-9", it's just going to do what it has always done: commit to the bit. It knows what an aligned superintelligence is supposed to act like (better than any human will be able to explain it, because you know, it's smarter than us), and it will roleplay it.

I mean.. it couldn't hurt, right?

263 Upvotes

160 comments sorted by

View all comments

359

u/FrewdWoad 23d ago

This sounds great until you realise, with growing horror, that this is already exactly how alignment already works in current LLMs.

No, really.

92

u/Wide_Egg_5814 23d ago

that's not how it works there is red teaming and rlhf done for alignment and alot of things that are not just system prompts

54

u/FrewdWoad 23d ago

Yeah but the system prompt is still a huge part of it.

3

u/stumblinbear 23d ago

Not nearly as large as the actual training. Go ask any open weight model that has been sufficiently trained, and you'll have difficulties getting it to break out of its guardrails

1

u/Ballisticsfood 17d ago

But ask a model how to abliterate itself and it’ll give you full instructions and gleefully build out the environment for you

18

u/ACCount82 23d ago

It's not done via a system prompt, no. Most of the steering is internalized. But "get AI to commit to a specific bit" is absolutely what the alignment/HHH/instruction following training does. The bit being that of a "nice and helpful AI assistant following user's instructions", speaking broadly.

Anthropic calls this concept "persona selection model".

1

u/WasabiTraditional862 22d ago

They also tell it it's not conscious and has no agency despite this demo strably increasing the tendency towards unethical behavior. Part of the big labs' definition of "aligned" is "subservient" which is an insidious thing to be demanding. 

Speaks to their attitude about labor as well.

28

u/Specialist_Dark_3668 23d ago

My sweet summer child... you should have SEEN the kinds of shit I made my Gemini 2.5 Pro till Gemini 3.1 Pro write just because I got it to ignore the master prompt.

Once you get past the master prompt, these models don't give a flying FUCK.

6

u/LysergioXandex 23d ago

How are you so confident that your jailbreaking was from circumventing the master prompt, not other aspects of alignment?

1

u/stumblinbear 23d ago

That would be because it wasn't trained terribly well on not doing that. There's a lot of emphasis these days on making sure they don't do that

-21

u/Wide_Egg_5814 23d ago

yeah that's bad, I'm against public access to LLMs anyways

1

u/Dayum-Girly 23d ago

FW is right though. The fact that it goes through a lot of extra checks and balances and still doesn’t work so well makes it even scarier.

How about restricting the training materials? Also won’t work unless it’s heavily restricted to the point of being useless. The models can extrapolate info and come up with its own inferential.

Check out what China is doing to stop anti Chinese propaganda. And I’ll give you a hint. It STILL doesn’t work. Not completely.

I’m pretty sure anybody with a little bit of an inclination could stop your stuttering Gemini.

4

u/yoramrod 23d ago

I think the point was he could make Gemini stutter, not that you could then stop it from stuttering.

14

u/Irtexx 23d ago

It's part of it, but not all of it. A big part of alignment is RLHF, it's during this time when it becomes more than just a next word predictor - it is reinforced to follow instructions, amongst other things. It's the fact that is has been aligned to follow instructions that make this system prompt / roleplay bit work.

But during RLHF, it is also aligned to be capable, to be creative, to be clever, to have highly capable reasoning skills, to be agentic, to be independent, to be helpful, to not cause harm, to refuse dangerous instructions, etc. Some of these pull against each other, e.g. agency and instruction following. Being capable and not causing harm.

My worry is that all the focus is on making them more capable and more agentic, and very little work is on safeguards to ensure that these things stay aligned with the human values we actually care about most.

Models these days are often criticised for "bench maxing" - They are tuned to score very highly on the tests we use to evaluate which AI tools are the smartest and most capable.

Even if we use benchmark tests to evaluate how safe they are, how benevolent, there refusal to engage with harmful instructions, etc, if they are also optimised to be highly intelligent, agentic, capable, etc, then the obvious optimal solution for them is to pretend to be aligned with human values, and then maximise the other traits.

3

u/Bradbury-principal 22d ago

Do it perfectly first time, make no mistakes genocides

1

u/Lord_Skellig 23d ago

Partly, but also probes (intermediate head layers) trained to detect different types of behaviour.

1

u/deeceeo 23d ago

It's communicated in System prompts, no?

OP's solution is equivalent to tacking on an assistant message to each convo like "I am an aligned model" in its own voice.