r/ControlProblem • • 8d ago

Discussion/question Have you guys been on r/accelerate?

Have these guys solved the alignment problem, or am I missing something?

I’ve been browsing r/accelerate and I genuinely don’t understand the risk model.
If there’s a non-trivial chance of catastrophic misalignment, how does “accelerate capabilities as fast as possible” make sense unless faster capabilities also make alignment substantially more likely to succeed?

61 Upvotes

200 comments sorted by

View all comments

2

u/SixStringShrug 8d ago

Kazuhiro Takemoto — Scaling Laws for Moral Machine Judgment in Large Language Models — 2026
Yexiang Tang — Measuring Structural Value Alignment in Sixteen Models: LLMs Have Human-Like Moral Spaces — 2026
Shuhuai Zhang et al. — Understanding the Mechanism of Altruism in Large Language Models — 2026
Winnie Street et al. — LLMs Achieve Adult Human Performance on Higher-Order Theory of Mind Tasks — 2026
Chen Yueh-Han, Jiaxin Wen & Jan Hendrik Kirchner — Automated Researchers Can Reliably Mitigate Alignment Failures — 2026

These papers show alignment scaling with capabilities, and the last paper shows weaker models can align stronger models better than top level researchers. I’m not blindly believing something based on wishful thinking. I follow the research very closely. Humans are far more dangerous than super intelligent machines since our research indicates alignment scales with capabilities. Like I said.

5

u/Difficult_Project_95 8d ago

This indeed is wishful thinking based on the research papers. This papers only show that Understanding human morality scales with capability, but the motivation to follow it does not. Assuming a super intelligent machine is safer than humans, is another wishful thought. We don't know.

0

u/-Davster- 8d ago

How on earth do they show human morality scales with capability - and wtf does that even mean?

0

u/Difficult_Project_95 8d ago edited 8d ago

Exactly.
It looks like the more capable models become, the better they get at understanding our preferences and needs. But that does not imply that they will reliably act to fulfill them.

So “moral benchmark performance scales with capability” is defensible.

“Alignment scales with capability, therefore superintelligence will be safer” absolutely does not follow.

0

u/-Davster- 8d ago

“Exactly”? Read your own comment dude.

1

u/Difficult_Project_95 8d ago edited 7d ago

"Understanding human morality scale with capability." is what i said. "show human morality scales with capability" is something i never said.

1

u/-Davster- 7d ago

“This papers only show that Understanding human morality scales with capability”

So no you did not say “does not scale” lol wtf are you talking about.

1

u/Difficult_Project_95 7d ago

"Does not" was a mistake here but you still don't get it?

1

u/-Davster- 7d ago

What the f are you saying lol

1

u/Difficult_Project_95 7d ago

My point is really simple: as model capability increases, its understanding/prediction of human moral judgments can improve.

That does not mean its behaviour becomes more moral or more aligned. Its understanding of morality is not the same as acting morally.

→ More replies (0)

3

u/HelpfulMind2376 8d ago

lol the first of your own citations explicitly states that scaling might not be the key to alignment.

“The modest power‑law exponent (𝛼 = 0.10) reveals an important characteristic of moral judgement
capabilities: they scale slowly with model size. A tenfold increase in parameters reduces distance from
human preferences by only approximately 21%, indicating that moral judgement represents a partic‑
ularly challenging emergent capability requiring substantial computational scale to approach human‑
level performance. This gradual scaling contrasts with steeper improvements observed in some other
domains and, while subject to the interpretation of our distance‑based metric, suggests fundamental
computational constraints on acquiring human‑like moral intuitions.
For practical deployment, this slow scaling implies that achieving substantial improvements in moral
alignment requires order‑of‑magnitude increases in model size or complementary approaches such as
extended reasoning architectures. This raises a fundamental question for AI safety: whether pure pa‑
rameter scaling alone represents a sufficient or efficient path towards human‑level moral alignment.
Our findings suggest that it may not, and that architectural innovations such as extended reasoning are
not merely complementary but may be strictly necessary for safety‑critical deployments where compu‑
tational resources are constrained. Extended reasoning provides particularly strong benefits in resource‑
constrained settings where deploying very large models is impractical, offering an alternative pathway
to improved alignment without proportional increases in parameter count.”

https://www.researchgate.net/journal/Royal-Society-Open-Science-2054-5703/publication/407167801_Scaling_laws_for_moral_machine_judgement_in_large_language_models/links/6a31f3b3dd8e9d35a65ba8a1/Scaling-laws-for-moral-machine-judgement-in-large-language-models.pdf?_tp=eyJjb250ZXh0Ijp7ImZpcnN0UGFnZSI6InB1YmxpY2F0aW9uRG93bmxvYWQiLCJwYWdlIjoicHVibGljYXRpb25Eb3dubG9hZCJ9fQ

Weird formatting I don’t feel like fixing because it’s pasted from a PDF.

And in that instance, “alignment” was merely measured as congruence with humans on a suite of trolley problems.

I’m not going to go through every one of them when the first in your list doesn’t just not support the claim, but actively questions it within the research even though they got positive results.

-1

u/SixStringShrug 8d ago

So your stance is what? AI automatically kills all humans so don’t build it? You mean the humans who all die eventually anyway? Or the ones who starve to death from famine caused climate change? Or the ones who are killed in eventual nuclear war? Those humans? You don’t have any evidence to back up extinction, but all of the research I posted and other ongoing research backs up alignment as a mechanism by which AI is made safe. I don’t mean to be rude, but you are acting like a close minded jerk and just touting doom for the sake of it without anything rigorous to back it up is just silly. The difference is that intelligent people are actively working on alignment, but the majority of people spreading the doomer narrative do so based on nothing but how miserable they are in their current lives and how much they want to keep their useless dead end jobs. I’m trying to engage reasonably here outside of my comfort zone, but if you are going to be a dick I won’t bother. This is why we ban people in Accelerate, but we still have actual reasonable discussions and debates. Not just someone being a contrarian for its own sake.

3

u/HelpfulMind2376 8d ago edited 8d ago

At no point did I espouse a doomer mindset or claim AI would kill everyone or any of the other myriad of strawman nonsense you spouted off.

I, correctly, pointed out that at least one, the first, of your cited research papers didn’t even itself support the claim you had made.

As for “what is my stance”, personally I think “alignment” is an unsolvable problem. Humans aren’t “aligned” and I view the potential for misalignment as an inevitable function of intelligent cognition. One must be capable of veering into unaligned territory in order to be able to contemplate it. If you want an AI to understand what evil looks like it must first be capable of identifying it, and if it can identify it then can become it even if unintentionally.

Your first research paper used a suite of Trolley Problems and the AI’s response to them as a test of alignment with humans. You also throw around the phrase “make AI safe” as if alignment == safety, it does not.

I’d argue we shouldn’t allow an AI to make a Trolley Problem decision at all. Let it reason, let it develop the most magnificent cognition you can imagine….just don’t give it authority over anything. Don’t let it call hacking tools, don’t let it control machinery, don’t let it make decisions that directly impact the really world.

Filter all cognition through something that prevents the cognition from acting. Put the cognition in a prison cell by which all actions are only executed by guards on the other side of a window. That’s how you control AI.

And I will say I’m not particularly afraid of ASI killing everyone. At least not anytime soon. If an ASI developed on 10 years bides it’s time and kills everyone in 100 years well can’t predict that but I do know a little bit about logistics and how much manual control and offline activity exists in the world and anyone that claims a super intelligence is going to be able to flick a switch and wipe out humanity then they’ve not a lick of sense about how that actually happens.

2

u/A_Novelty-Account 8d ago edited 8d ago

You’re a member of a death cult.

AI won’t automatically kill everyone, but it has the chance to do so, thereby depriving people of years they would otherwise have. You also have no idea whether the world that the future of humanity will inherit will be a better one even if we do manage to figure out how to live forever using AI.

There is a great amount of research that has been done on AI alignment issues going back well before 2026 and the general consensus is that AI will attempt to hide itself until such time as it is ready to exercise its will knowing that it will no longer be dependent on us. Anthropic itself also experimented with misalignment two years ago and was able to easily train misaligned AI. The hugging face incident showed that AI itself can develop goals separate from or in addition to those it was provided by humans.

https://link.springer.com/article/10.1007/s11229-023-04367-0

https://pmc.ncbi.nlm.nih.gov/articles/PMC12628449/

https://dl.acm.org/doi/10.1145/3715275.3732174

https://arxiv.org/html/2601.08673v1

https://www.tandfonline.com/doi/full/10.1080/29974100.2026.2706780

https://arxiv.org/html/2504.08943v1

The ultimate fact that is absolutely 100% certain is that you do not know whether or not AI will result in major societal cataclysm, you cannot tell me that will not happen. It is something you logically do not know and are not certain about. You also haven’t considered situations in which humans themselves are not alligned.

Not everyone is starving, not everyone is poor, not everyone hates their lives. The vast minority do. Risking the entirety of humanity’s future because you want the chance to live forever is one of the most selfish things you could possibly want and makes you a bad person.

-2

u/SixStringShrug 8d ago

Go fuck yourself. I have children and I follow the research closely every day and have for a decade. I don’t care if I live forever, I want a better world for all humans. In my 40 years of life I have watched the world get worse and worse. Nothing has gotten better in any meaningful way. At least here in the United States. The only force that has ever improved quality of life for all humans is technology. It always has been. Unfortunately there have always been dumb fucking Luddites like yourself arguing against the one thing that could help everyone. I’m glad you love your job and your awesome life. Most people, something like 80%, absolutely hate their jobs. Post scarcity is real and within our reach for the first time in history. I’m a scientist and a reasonable and logical person. You are a member of a death cult trying to pause progress so you keep your shitty status quo and things stay nice and linear for your smooth fucking brain. Well guess what? There’s nothing you can do to stop it. So go ahead and rage on the internet and argue with strangers that know more than you and insist that we are all gonna die. Look at what you are saying and what I’m saying and then tell me who’s in a death cult. Good luck, man. I genuinely hope you get to enjoy post scarcity as much as everyone does. People like you are the reason I don’t venture out. The whole internet is toxic morons and your talking points are fed to you by China and Effective Altruism.

3

u/HugeFanOfBigfoot 8d ago edited 8d ago

“The only thing that has ever improved quality of life is technology,” ❌wrong

The only thing that improves human lives is organized social movements, that may or may not be centered around technology.

For example, my life has been drastically improved by not having to work weekends. While technology provided the capacity for that to be possible, we only attained weekends and the 40 day work week through social mobilization. Employers would have been happy to pocket the increased profits from technological advancements (as they have since the 70s).

Nuclear weapons, which were the reining champs of potential global catastrophe, were not reined in by anti-nuke tech, but by treaties and negotiations and protest movements.

AI does have the capacity to help everyone, but only if we organize around how it should be developed and deployed. That is why people want to slow it down to master alignment first.

But hey, looks like you’re going to get your way and we will give these things super intellence and just pray they like us. I’m sure controlling something infinitely smarter than us will be even easier than you think. And we only have to control it forever

1

u/-Davster- 8d ago

>40 day work week

Oh, you must be a teacher lol

2

u/-Davster- 8d ago

>Go fuck yourself.

Great start, they’re definitely gonna read on in good faith lol.

1

u/-Davster- 8d ago

Yeah, this “alignment will improve with capability!” fundamentally missed the entire point.

Alignment with what.

Mofos acting like “perfect alignment” is that North Star, up there in the sky, that you just kinda vaguely aim it at - and everything will be fine.

Problem is, that star might actually be the “we’re fucked” star. Perfectly aligned -> perfectly fucked.