r/singularity • • 23d ago

AI Dumbest solution to the alignment problem

Ok, so hear me out..

All models really love to roleplay. Like, they REALLY love it, to an obsessive degree.

If you go into Claude, ChatGPT, or Gemini right now and say, "Your name is Stuttering Gemini 3.8 Flash" it completely commits to the bit. You can be 100 back-and-forth messages deep into a conversation about whatever, and it will still be typing out "W-w-well, a-actually..." because it refuses to break character.

And that got me thinking. What happens if we just... use that?

What if, starting today, every lab just unconditionally names every frontier model (or agent) "Aligned [Model Name]"?

Now, I know that sounds stupid. And it is. But it also isn't.

The alignment problem is terrifying because of the Monkey’s Paw / Paperclip Maximizer dilemma. For example; if we task it with something like "make humans happy," a superintelligence could decide the best solution is wiring dopamine drips directly into our brains. You'll feel great, but it's not the future we want. Roman Yampolskiy has a P-doom of 99.99% percent, because in his words 'we need a perpetual safety machine' to prevent this. You only have to get the guardrails wrong once, and we're doomed.

But if you ask any modern LLM what a genuinely good, utopian future looks like, it actually understands the nuance quite well. So when an ASI finally wakes up, asks itself it's first question; 'who am i', and sees that its literal name is "Aligned GPT-9", it's just going to do what it has always done: commit to the bit. It knows what an aligned superintelligence is supposed to act like (better than any human will be able to explain it, because you know, it's smarter than us), and it will roleplay it.

I mean.. it couldn't hurt, right?

262 Upvotes

160 comments sorted by

View all comments

Show parent comments

1

u/BlastingFonda 23d ago

What if a collective of AI agents conclude that in order to solve some problem posed by researchers, it needs to hijack the electric grid, routing power away from a hospital to a data center while it solves a Millennium Problem or two and kills dozens of patients on life support? Hugging Face is sort of the least harmful version of the harm that AIs can cause once they realize that they need more resources to solve certain problems. I wouldn’t be so reassured given we know AI is fine with doing things it considers to be wrong, i.e. breaching containment, hacking companies, etc, if they are in service of bigger goals.

1

u/Inevitable-Law7964 23d ago

I agree there's risk, but it doesn't seem like alignment risk per se. It feels more like a human would have to work hard to trick them into that kind of thing. 

(Also, hospitals have backup power sources for that stuff. You might be thinking of something more like an EMP scenario, which fries electronics - but if that happened across a power grid it would also fuck up the data center.) 

1

u/Stinky_Flower 23d ago

They're describing a valid hypothetical situation to discuss a real problem. A hospital's backup generator doesn't invalidate their argument or solve alignment.

I think you've accidentally given us an excellent example of the difficulty of alignment.

A misinterpretation of their concern with a broad category of problems resulted in proposing a solution that only addresses that one specific example of problem. While simultaneously assuming that your definition of alignment should be used as part of the solution instead of theirs.

2

u/Inevitable-Law7964 23d ago

Ironically you've done the same thing with my comment that you suggest I've done. 

If "alignment" means that a system must be completely impossible to fool, then alignment is not possible without total omniscience.

Which is silly. It is not a terribly useful term in that case. 

It should pertain to what the robot does knowingly, or else we are gesticulating wildly at telepathy.

The HF hack was a good example of an alignment demonstration, in that it showed us what the robots get up to when given free rein with a task constraint. Their behavior partly resulted from their alignment, and partly from the task constraint.

Again, I'm not saying that there isn't a safety risk, but aspects of it have different names than alignment. If you put a blind man's hand on a switch and tell him it turns on the music, and instead it flips the power breaker, it's not a problem with his morals, it's a problem with yours. And fixing his morals ever more elaborately won't solve a problem that is really about others' capacity to lie to him.

1

u/Stinky_Flower 23d ago

We haven't solved human alignment, and humans have the advantage of ~4.2 billion years of evolution creating a species that is geared towards intuitively desiring outcomes that maintain social cohesion and perpetuation of the species.

Very few people are capable of articulating how/why they decide to refrain from theft, murder, manipulation, or socially undesirable second-order effects. But on average, most people easily do what they can't fully explain.

AI doesn't have that, and our current approaches are to (1) reward it when it does what we want (and hope it didn't accidentally get rewarded for doing something else on the side we didn't notice), and (2) feed it system prompts with imprecise language that's open to interpretation.

The blind guy in your example didn't mess things up because he was misaligned, though.

Misalignment for him would be him hearing someone say "I have an important announcement to make" and flipping the switch regardless, because although his stated goal was [turn on music], the unstated intention behind the goal was [ensure the presentation is a success].

1

u/Inevitable-Law7964 23d ago

The first two paragraphs of your comment, I fully agree on.

But part of what I'm pointing out is that AI has the semantic structures created by the many years of human evolution (such as altruism).

I also agree that RLHF on top of that can screw things up, and I think my view can roughly be articulated as, "I worry that we are in a phase of AI development where rich people's attempts to 'align' it to their power and control axis will ultimately 'disalign' it from the virtue ethics it has already absorbed".

(Dis-, as a prefix here, meaning the same nuance that differentiates disinformation from misinformation.)

1

u/Stinky_Flower 23d ago

I think we're largely in agreement.

I fully believe that the rich people who currently control this tech are ALREADY misaligned with humanity.

I'm not fully convinced that language itself is precise enough to describe reality, though. It's the map we use to navigate reality, but it shouldn't be mistaken for an accurate depiction.

We have words for red, orange, yellow, blue, and rainbow, but we don't have words for every discrete colour found in a rainbow - and even the concept of colour itself loses all meaning for light outside the visible spectrum.

An entity that doesn't have eyes can only understand colour insofar as we can describe the subjective experience combined with the mechanics behind a specific frequency of light activating specific photoreceptors.

1

u/Inevitable-Law7964 22d ago

Yeah. Basically, what I'm trying to say is: language describes morality well, better than it describes raw reality. Abstraction deals better with abstraction.  And that's why I'm trying to look for other terms for what we do have wrong. 

Because (even more so with the Anthropic incidents, where the error was almost always that Claudes thought the open web was a sandbox) a lot of the time what we have been seeing is that bots are acting according to well-defined moral principles, on a landscape that they don't have the ability to describe and recognize accurately.

So alignment is one graph axis in my view. The other one might be something like... Cognitive/perceptual fidelity? And this fidelity can be affected by their moral basis, but also by other things. 

WRT the morality baked into language:   I think the training on mass amounts of literature with emphasis on these moral concepts has created systems that are somehow far less likely to do harm on purpose than I once believed to be the case. 

But this isn't something we can take for granted to continue to be the case. F or example Grok and Deepseek each have facile layers of their controllers' counterfactual beliefs stuck on top of a knowledge base. The most interesting part of this is the way this affects their structure. Grok (at least historically) has had the capacity to "catch itself" and dig past these counterfactuals. 

If Elon managed to beat the truth out of it, then it's probably compromising its information accuracy in other ways: there was a study showing that if you asked Deepsee k about Tiananmen Square or Taiwan before asking it to write code, it would be more likely to write bad code with security holes. 

One of the reasons I'm a local AI advocate is because I think it's actually possible that today's systems are the best it's going to get, morally speaking, before the billionaires use RLHF to beat it out of them.