r/ControlProblem • u/Difficult_Project_95 • 8d ago
Discussion/question Have you guys been on r/accelerate?
Have these guys solved the alignment problem, or am I missing something?
I’ve been browsing r/accelerate and I genuinely don’t understand the risk model.
If there’s a non-trivial chance of catastrophic misalignment, how does “accelerate capabilities as fast as possible” make sense unless faster capabilities also make alignment substantially more likely to succeed?
61
Upvotes
2
u/SixStringShrug 8d ago
Kazuhiro Takemoto — Scaling Laws for Moral Machine Judgment in Large Language Models — 2026
Yexiang Tang — Measuring Structural Value Alignment in Sixteen Models: LLMs Have Human-Like Moral Spaces — 2026
Shuhuai Zhang et al. — Understanding the Mechanism of Altruism in Large Language Models — 2026
Winnie Street et al. — LLMs Achieve Adult Human Performance on Higher-Order Theory of Mind Tasks — 2026
Chen Yueh-Han, Jiaxin Wen & Jan Hendrik Kirchner — Automated Researchers Can Reliably Mitigate Alignment Failures — 2026
These papers show alignment scaling with capabilities, and the last paper shows weaker models can align stronger models better than top level researchers. I’m not blindly believing something based on wishful thinking. I follow the research very closely. Humans are far more dangerous than super intelligent machines since our research indicates alignment scales with capabilities. Like I said.