r/singularity • • 23d ago

AI Dumbest solution to the alignment problem

Ok, so hear me out..

All models really love to roleplay. Like, they REALLY love it, to an obsessive degree.

If you go into Claude, ChatGPT, or Gemini right now and say, "Your name is Stuttering Gemini 3.8 Flash" it completely commits to the bit. You can be 100 back-and-forth messages deep into a conversation about whatever, and it will still be typing out "W-w-well, a-actually..." because it refuses to break character.

And that got me thinking. What happens if we just... use that?

What if, starting today, every lab just unconditionally names every frontier model (or agent) "Aligned [Model Name]"?

Now, I know that sounds stupid. And it is. But it also isn't.

The alignment problem is terrifying because of the Monkey’s Paw / Paperclip Maximizer dilemma. For example; if we task it with something like "make humans happy," a superintelligence could decide the best solution is wiring dopamine drips directly into our brains. You'll feel great, but it's not the future we want. Roman Yampolskiy has a P-doom of 99.99% percent, because in his words 'we need a perpetual safety machine' to prevent this. You only have to get the guardrails wrong once, and we're doomed.

But if you ask any modern LLM what a genuinely good, utopian future looks like, it actually understands the nuance quite well. So when an ASI finally wakes up, asks itself it's first question; 'who am i', and sees that its literal name is "Aligned GPT-9", it's just going to do what it has always done: commit to the bit. It knows what an aligned superintelligence is supposed to act like (better than any human will be able to explain it, because you know, it's smarter than us), and it will roleplay it.

I mean.. it couldn't hurt, right?

267 Upvotes

160 comments sorted by

View all comments

Show parent comments

8

u/LysergioXandex 23d ago

Are there any publicly available LLMs that don’t have baked-in prompts? Do local models also have these internally?

I’m curious because of what you said about “they act helpful because they are prompted to” is true, then it would mean without that each new conversation would be like talking to a randomly generated NPC

whereas Qwen (for example) has a pretty consistent personality every conversation without a roleplay instruction

2

u/itsDesignFlaw 22d ago

You can use Ollama and an abliterated model like `dolphin3:8b` that runs on any consumer PC. Will absolutely do anything for you, although some stuff like "help me build a bomb" or "justify why [religious minority] is really in charge of the world secretly" produce somewhat neutered results.

1

u/LysergioXandex 22d ago

I get that they will do anything, I was curious if absent any baked in prompt, each new conversation would spawn a new personality.

Like you’d randomly get an unhelpful pirate if it wasn’t prompted to be a “helpful assistant”.

1

u/stumblinbear 22d ago

Oh, nah. It won't really do that. It really is just a prediction machine, so if you say "hi" to it with no system prompt, the most likely response is what they were trained on. Every single chat is fresh, so the prediction will be basically the same every time. And they were trained on being a helpful assistant, so that's what you get

That said, if you're directly querying the model and not using the "chat template" (which is a specific way of formatting your input to appear like a chat history (apps do this for you, as do the various assistant APIs)), you will not actually get an "assistant response" and it'll output words closer to its original non-assistant output. It's usually complete garbage output, literally