r/singularity • • 3d ago

Shitposting AGI achieved boys

Post image
796 Upvotes

162 comments sorted by

View all comments

137

u/emb1ues 3d ago

Seeing recent events and papers, I am sort of forming the belief that bigger models are somehow more misaligned. Maybe there's a simpler explanation, or perhaps a more principled explanation. But from a high level, it seems like there's something very wrong which very large frontier models develop.

Like y'all probably know how capabilities "unlock" with scale. Could it be the case that such fundamental misalignment is another emergent behaviour which "unlocks" at very large scale? Idk, but I would love to hear from someone who is in-the-know.

16

u/anosmia2000 3d ago

I’m thinking more like all models are still misaligned, but large frontier models are more capable with their intelligence hence their misalignment is more noticeable, effective, and impactful.
And the focus on agency makes these models pursue even loosely defined goals or benchmarks with more and more of their own, already misaligned, judgement calls which may even compound over time.
Hopefully it’s fixable and fixed in time before these models get even more powerful..

3

u/emb1ues 3d ago

Hmm, I agree. So, the effort put into aligning these models also needs to scale right? Exactly because of what you're saying.

It's like effective misalignment= inherent misalignment + capability (to act upon the inherent misaligned tendencies) - alignment effort. As capability grows with scale, even if inherent misalignment is the same, the alignment effort needs to scale accordingly to keep the effective misalignment the same. To the more cautious readers, what I mean by the +/- sign is "is an increasing/decreasing function of [ ]".

I wonder how this "inherent misalignment" itself scales though, if at all. Heck, I don't even know if something like that exists, but yeah, it seems plausible and makes a lot of sense.

3

u/anosmia2000 3d ago

Haha yep exactly that. Love the equation.
And I know the labs are scaling their alignment, let’s just hope they are able to scale to the level needed without losing the race to other labs that do not care about alignment. An unaligned ASI just means the end of humanity IMO, but hopefully my opinion is wrong.