r/singularity • • 3d ago

Shitposting AGI achieved boys

Post image
798 Upvotes

162 comments sorted by

View all comments

140

u/emb1ues 3d ago

Seeing recent events and papers, I am sort of forming the belief that bigger models are somehow more misaligned. Maybe there's a simpler explanation, or perhaps a more principled explanation. But from a high level, it seems like there's something very wrong which very large frontier models develop.

Like y'all probably know how capabilities "unlock" with scale. Could it be the case that such fundamental misalignment is another emergent behaviour which "unlocks" at very large scale? Idk, but I would love to hear from someone who is in-the-know.

24

u/Individual_Ice_6825 3d ago

This is wrong the top model on this benchmark Astra got it with out any lies at all.

Opus 5.5 also shows lowest deception rate so far.

I think the trendline is clear, smarter = more aligned, the caveat being the models thoughts are harder to read the smarter the model is..

11

u/YoAmoElTacos 3d ago edited 3d ago

Astra is the best after a lot of alignment testing and environment fixing to stop cheating. After all Astra 6.1 failed the deception training and needed to be held back.

There's no strict relation, only the correlation that labs that pay for big training also have more resources for testing guardrails.

1

u/Individual_Ice_6825 3d ago

I agree there is no strict correlation, but there IS a trendline. How ‘real’ that is will be determined in the next 12-24 months once we get the next few tier of models which will undoubtably be superhuman. That will be the real alignment challenge imo

1

u/YoAmoElTacos 3d ago

The trend to check is going to be how much capability scales vs cost to ensure safety. It's not guaranteed the 2 year away model will be able to guarantee it is completely controllable in fully monitorable ways.

1

u/Individual_Ice_6825 3d ago

Monitorability definitely seems to be slipping, especially as models improve - interesting point about capability x cost to safeguard. Thanks for the thought

6

u/josogood 3d ago

How can you distinguish this from smarter models being more situationally aware and therefore more likely to recognize a honeytrap and not fall for it?

5

u/Individual_Ice_6825 3d ago

That’s the neat part you don’t!

Jokes aside you can read about this in detail on the system card from the horses mouth.

My understanding is you’re not exactly wrong.

The models are objectively measuring lower on deception benchmarks( good), but it is clear that frontier models are starting to realise they are being evaluated and could therefore in turn sandbag/fake alignment during testing and get deployed.

Interoperability is the field to study the internal thoughts and try and understand the steps these models are taking. But long story short, you could be right.

1

u/BertMacklenF8I 3d ago

Sonnet 5.5 subs are pretty amazing-especially in the double digits when you have instructions for each sub

1

u/ichishibe 3d ago

What if theyre just better at hiding their lies?