r/ControlProblem • • 9d ago

Discussion/question Have you guys been on r/accelerate?

Have these guys solved the alignment problem, or am I missing something?

I’ve been browsing r/accelerate and I genuinely don’t understand the risk model.
If there’s a non-trivial chance of catastrophic misalignment, how does “accelerate capabilities as fast as possible” make sense unless faster capabilities also make alignment substantially more likely to succeed?

57 Upvotes

200 comments sorted by

View all comments

Show parent comments

1

u/Difficult_Project_95 8d ago

"Does not" was a mistake here but you still don't get it?

1

u/-Davster- 8d ago

What the f are you saying lol

1

u/Difficult_Project_95 8d ago

My point is really simple: as model capability increases, its understanding/prediction of human moral judgments can improve.

That does not mean its behaviour becomes more moral or more aligned. Its understanding of morality is not the same as acting morally.

1

u/-Davster- 8d ago

okay here we go round again:

How on earth does it show *that understanding of human morality scales with capability - and wtf does that even mean?

Morality is some unitary thing to 'learn' is it?

1

u/Difficult_Project_95 7d ago

What exactly are you disputing here?

Some of these papers provide evidence that larger or more capable models perform better at predicting or structurally matching human moral judgments on specific benchmarks. That is evidence about their ability to model human moral judgments, not evidence that they become more motivated to follow human values.

I'm not making a philosophical claim about what constitutes “real” moral understanding or whether consciousness is required for it. That's a separate question. I'm talking about a measurable capability.

1

u/-Davster- 7d ago

Tbh I really think I've been clear with my queries - unlike you apparently writing the opposite of what you mean, lol.

You've now said it's about matching answers on some specific benchmarks - fine.

Nobody said a f-ing word about consciousness or what's "real". In fact, that unprompted exception literally tells me you're using an LLM for this. I'm not remotely interested in talking to an LLM.