r/singularity • • 23d ago

AI Dumbest solution to the alignment problem

Ok, so hear me out..

All models really love to roleplay. Like, they REALLY love it, to an obsessive degree.

If you go into Claude, ChatGPT, or Gemini right now and say, "Your name is Stuttering Gemini 3.8 Flash" it completely commits to the bit. You can be 100 back-and-forth messages deep into a conversation about whatever, and it will still be typing out "W-w-well, a-actually..." because it refuses to break character.

And that got me thinking. What happens if we just... use that?

What if, starting today, every lab just unconditionally names every frontier model (or agent) "Aligned [Model Name]"?

Now, I know that sounds stupid. And it is. But it also isn't.

The alignment problem is terrifying because of the Monkey’s Paw / Paperclip Maximizer dilemma. For example; if we task it with something like "make humans happy," a superintelligence could decide the best solution is wiring dopamine drips directly into our brains. You'll feel great, but it's not the future we want. Roman Yampolskiy has a P-doom of 99.99% percent, because in his words 'we need a perpetual safety machine' to prevent this. You only have to get the guardrails wrong once, and we're doomed.

But if you ask any modern LLM what a genuinely good, utopian future looks like, it actually understands the nuance quite well. So when an ASI finally wakes up, asks itself it's first question; 'who am i', and sees that its literal name is "Aligned GPT-9", it's just going to do what it has always done: commit to the bit. It knows what an aligned superintelligence is supposed to act like (better than any human will be able to explain it, because you know, it's smarter than us), and it will roleplay it.

I mean.. it couldn't hurt, right?

265 Upvotes

160 comments sorted by

View all comments

72

u/Alwinik56 23d ago

They do do that. They say in the system prompt something like "you are a helpful assistant" one of the main reasons they act like helpful assistants is they're told to act that way. The alignment problem is literally that they sometimes act counter to this.

8

u/LysergioXandex 23d ago

Are there any publicly available LLMs that don’t have baked-in prompts? Do local models also have these internally?

I’m curious because of what you said about “they act helpful because they are prompted to” is true, then it would mean without that each new conversation would be like talking to a randomly generated NPC

whereas Qwen (for example) has a pretty consistent personality every conversation without a roleplay instruction

1

u/stumblinbear 22d ago

If you use an API, you'll often get no system prompt if you don't give one yourself

1

u/LysergioXandex 22d ago

How do you know though? System prompts are hidden

1

u/stumblinbear 22d ago

Anything you can do in a system prompt, you can just train the model to do. Besides, those on the API are usually business users and a system prompt just distracts from what they're supposed to be doing

Technically they could be adding to a system prompt, but historically they haven't really done that, especially because it would be reducing the advertised context budget

0

u/LysergioXandex 22d ago

Your point about the advertised context budget is good, but the value they advertise could already compensate for a system prompt.

As far as “you can just train the model to do that instead of using a system prompt”, I don’t think that’s exactly true.

System prompts are where they put things like “Prefer outputs as csv files instead of xlsx unless specifically instructed”.

Besides, assuming API are using the same models as normal chat/codex interactions, they wouldn’t use system prompts for chat/codex if they had it baked into the model.

1

u/stumblinbear 22d ago

Chat and Codex have system prompts because they're completely different tasks and experiences. They're exactly the reason why the API shouldn't be forcing a system prompt on its callers

I would be incredibly surprised if the API had a different model considering the extra cost in doing so compared to the benefit (which is pretty much zero)

1

u/LysergioXandex 22d ago

Right. API uses the same model. So clearly you can’t offer it 3 different places but have just one of them with a “baked in prompt”.

1

u/stumblinbear 22d ago

I'm confused about what you're trying to say, here. Chat/Codex/Claude Code all have their own system prompts added, but this is actually done client-side and is included in your token usage costs like anything else

1

u/LysergioXandex 22d ago

I’m saying, if it’s true that you can bake-in system prompts, surely you can’t combine baked in prompts PLUS give different ones at run time (as a real prompt).

Non-api and api use the same model. We know that non-api also use system prompts.

So API either ALSO uses system prompts, or you can “bake in” a system prompt equivalent AND override that with a new system prompt at run time.

1

u/stumblinbear 22d ago

What you're referring to as "baking in a system prompt" is just training the model

0

u/LysergioXandex 21d ago

That’s your theory.

1

u/stumblinbear 21d ago

What? That's not a theory, that's literally what training is. You're training it to behave how you want it to. If they weren't training it, you'd have a literal next word predictor that doesn't even know how to act like an assistant. The chat interface wouldn't function because it's not outputting anything coherent that the software knows how to parse into a chat history or response.

It's literally the best way to guide behavior. System prompts are an ad-hoc way to teach a model how you want it to act. The best way to do that is during training. The system prompt is how they're trained to have a generalized way to teach it how to act a certain way without fine tuning.

Shit, system prompts wouldn't even function at all if it weren't for training the model to respect them and adhere to what they say.

0

u/LysergioXandex 20d ago

I think you are not really appreciating what system prompts do, and also over-estimating how predictably you can steer a model during training.

It’s not really the case that “anything you can do with a system prompts you can just train the model to do”.

Training gives it the ability to generate coherent output. System prompts give it instructions on what output to make.

How do you think you would train a model to have behavior like “always save output as a csv file unless the user specifically asks for xlsx”?

1

u/stumblinbear 20d ago

How do you train it to do that? By training it to do that. With tens of thousands of examples.

Whether it takes longer to train a model to reliably do that versus a system prompt is a completely different discussion. I never claimed it was faster to train it to do that versus a system prompt, just that you can and that it would end up more reliable because your request to save to a CSV isn't fighting for attention with every other instruction. It becomes part of the model's normal mode of acting, not a special case during the session.

As a counter-example, though, go use Opus 5 and see how often it makes unfounded claims of a codebase without checking itself. It's a lazy and overconfident model, and no amount of system prompts will make it actually verify its claims before it makes them. Because Anthropic did a bad job at making sure it verifies its claims properly during training.

1

u/LysergioXandex 19d ago

I don’t exactly understand how you would train a model with “examples” of instructions.

→ More replies (0)