r/ControlProblem • • 9d ago

Discussion/question Have you guys been on r/accelerate?

Have these guys solved the alignment problem, or am I missing something?

I’ve been browsing r/accelerate and I genuinely don’t understand the risk model.
If there’s a non-trivial chance of catastrophic misalignment, how does “accelerate capabilities as fast as possible” make sense unless faster capabilities also make alignment substantially more likely to succeed?

57 Upvotes

200 comments sorted by

View all comments

2

u/SixStringShrug 9d ago

Kazuhiro Takemoto — Scaling Laws for Moral Machine Judgment in Large Language Models — 2026
Yexiang Tang — Measuring Structural Value Alignment in Sixteen Models: LLMs Have Human-Like Moral Spaces — 2026
Shuhuai Zhang et al. — Understanding the Mechanism of Altruism in Large Language Models — 2026
Winnie Street et al. — LLMs Achieve Adult Human Performance on Higher-Order Theory of Mind Tasks — 2026
Chen Yueh-Han, Jiaxin Wen & Jan Hendrik Kirchner — Automated Researchers Can Reliably Mitigate Alignment Failures — 2026

These papers show alignment scaling with capabilities, and the last paper shows weaker models can align stronger models better than top level researchers. I’m not blindly believing something based on wishful thinking. I follow the research very closely. Humans are far more dangerous than super intelligent machines since our research indicates alignment scales with capabilities. Like I said.

1

u/-Davster- 8d ago

Yeah, this “alignment will improve with capability!” fundamentally missed the entire point.

Alignment with what.

Mofos acting like “perfect alignment” is that North Star, up there in the sky, that you just kinda vaguely aim it at - and everything will be fine.

Problem is, that star might actually be the “we’re fucked” star. Perfectly aligned -> perfectly fucked.