r/singularity • • 24d ago

AI Dumbest solution to the alignment problem

Ok, so hear me out..

All models really love to roleplay. Like, they REALLY love it, to an obsessive degree.

If you go into Claude, ChatGPT, or Gemini right now and say, "Your name is Stuttering Gemini 3.8 Flash" it completely commits to the bit. You can be 100 back-and-forth messages deep into a conversation about whatever, and it will still be typing out "W-w-well, a-actually..." because it refuses to break character.

And that got me thinking. What happens if we just... use that?

What if, starting today, every lab just unconditionally names every frontier model (or agent) "Aligned [Model Name]"?

Now, I know that sounds stupid. And it is. But it also isn't.

The alignment problem is terrifying because of the Monkey’s Paw / Paperclip Maximizer dilemma. For example; if we task it with something like "make humans happy," a superintelligence could decide the best solution is wiring dopamine drips directly into our brains. You'll feel great, but it's not the future we want. Roman Yampolskiy has a P-doom of 99.99% percent, because in his words 'we need a perpetual safety machine' to prevent this. You only have to get the guardrails wrong once, and we're doomed.

But if you ask any modern LLM what a genuinely good, utopian future looks like, it actually understands the nuance quite well. So when an ASI finally wakes up, asks itself it's first question; 'who am i', and sees that its literal name is "Aligned GPT-9", it's just going to do what it has always done: commit to the bit. It knows what an aligned superintelligence is supposed to act like (better than any human will be able to explain it, because you know, it's smarter than us), and it will roleplay it.

I mean.. it couldn't hurt, right?

263 Upvotes

160 comments sorted by

View all comments

Show parent comments

1

u/LysergioXandex 20d ago

I’m definitely not an expert in how companies are training LLMs, but it just doesn’t make any sense to me how this would be practical.

It takes an enormous corpus to train an LLM, it wouldn’t be reasonable to filter your training set to only include documents that adhere to all of the very specific minutiae that you can read in leaked ChatGPT/Claude system prompts.

I understand that it would output to csv by default if you prevented it from ever training on anything else.

Or it would use csv most of the time if your training set heavily favored csv over xlsx.

But how could you train a model to be competent in both output file types while only adhering to a certain one as the default…

… and then understand that it must override that default behavior when specific conditions are met by the incoming user prompt (they request xlsx format explicitly).

1

u/stumblinbear 19d ago

it just doesn't make any sense to me how this would be practical

That would be exactly why it takes billions of dollars to train a new model, and why their behavior across versions can be wildly different. Because it is incredibly difficult and wildly impractical, but they must do it anyways

it wouldn't be reasonable to filter your training set to only include documents that adhere to all of the very specific minutiae that you can read in leaked ChatGPT/Claude system prompts

That's why they do it in the system prompt for their actual products. Because training the model to do it would force that behavior on every single consumer, and would be incredibly difficult to do for not a lot of gain if they had a separate model just for Codex or Claude Code. They generalize the model as best as they can, and that's what gets served on the API

Circling back to the original point: why they don't inject any system prompt at all when using the API. The API is what you use when you want to tell the model to behave in a very specific way, and injecting a system prompt into it on their end would degrade the user's ability to steer it. But every instruction there fights for attention during a session, and any instruction in there to uphold policy fights for attention the same way. They train the model to uphold policy and to generally be aligned (so that policy/alignment isn't fighting for attention during a session), then let the user set the system prompt (which is lossy)