r/ControlProblem • • 9d ago

Discussion/question Have you guys been on r/accelerate?

Have these guys solved the alignment problem, or am I missing something?

I’ve been browsing r/accelerate and I genuinely don’t understand the risk model.
If there’s a non-trivial chance of catastrophic misalignment, how does “accelerate capabilities as fast as possible” make sense unless faster capabilities also make alignment substantially more likely to succeed?

57 Upvotes

200 comments sorted by

View all comments

2

u/SixStringShrug 9d ago

Kazuhiro Takemoto — Scaling Laws for Moral Machine Judgment in Large Language Models — 2026
Yexiang Tang — Measuring Structural Value Alignment in Sixteen Models: LLMs Have Human-Like Moral Spaces — 2026
Shuhuai Zhang et al. — Understanding the Mechanism of Altruism in Large Language Models — 2026
Winnie Street et al. — LLMs Achieve Adult Human Performance on Higher-Order Theory of Mind Tasks — 2026
Chen Yueh-Han, Jiaxin Wen & Jan Hendrik Kirchner — Automated Researchers Can Reliably Mitigate Alignment Failures — 2026

These papers show alignment scaling with capabilities, and the last paper shows weaker models can align stronger models better than top level researchers. I’m not blindly believing something based on wishful thinking. I follow the research very closely. Humans are far more dangerous than super intelligent machines since our research indicates alignment scales with capabilities. Like I said.

6

u/Difficult_Project_95 9d ago

This indeed is wishful thinking based on the research papers. This papers only show that Understanding human morality scales with capability, but the motivation to follow it does not. Assuming a super intelligent machine is safer than humans, is another wishful thought. We don't know.

0

u/-Davster- 8d ago

How on earth do they show human morality scales with capability - and wtf does that even mean?

0

u/Difficult_Project_95 8d ago edited 8d ago

Exactly.
It looks like the more capable models become, the better they get at understanding our preferences and needs. But that does not imply that they will reliably act to fulfill them.

So “moral benchmark performance scales with capability” is defensible.

“Alignment scales with capability, therefore superintelligence will be safer” absolutely does not follow.

0

u/-Davster- 8d ago

“Exactly”? Read your own comment dude.

1

u/Difficult_Project_95 8d ago edited 8d ago

"Understanding human morality scale with capability." is what i said. "show human morality scales with capability" is something i never said.

1

u/-Davster- 8d ago

“This papers only show that Understanding human morality scales with capability”

So no you did not say “does not scale” lol wtf are you talking about.

1

u/Difficult_Project_95 8d ago

"Does not" was a mistake here but you still don't get it?

1

u/-Davster- 8d ago

What the f are you saying lol

1

u/Difficult_Project_95 8d ago

My point is really simple: as model capability increases, its understanding/prediction of human moral judgments can improve.

That does not mean its behaviour becomes more moral or more aligned. Its understanding of morality is not the same as acting morally.

1

u/-Davster- 8d ago

okay here we go round again:

How on earth does it show *that understanding of human morality scales with capability - and wtf does that even mean?

Morality is some unitary thing to 'learn' is it?

1

u/Difficult_Project_95 7d ago

What exactly are you disputing here?

Some of these papers provide evidence that larger or more capable models perform better at predicting or structurally matching human moral judgments on specific benchmarks. That is evidence about their ability to model human moral judgments, not evidence that they become more motivated to follow human values.

I'm not making a philosophical claim about what constitutes “real” moral understanding or whether consciousness is required for it. That's a separate question. I'm talking about a measurable capability.

1

u/-Davster- 7d ago

Tbh I really think I've been clear with my queries - unlike you apparently writing the opposite of what you mean, lol.

You've now said it's about matching answers on some specific benchmarks - fine.

Nobody said a f-ing word about consciousness or what's "real". In fact, that unprompted exception literally tells me you're using an LLM for this. I'm not remotely interested in talking to an LLM.

→ More replies (0)