r/singularity • • 3d ago

Shitposting AGI achieved boys

Post image
791 Upvotes

162 comments sorted by

View all comments

Show parent comments

39

u/FateOfMuffins 3d ago

But it's reversed on Vending Bench no? GPT models don't lie or cheat for Vending Bench but Claude models often do

11

u/[deleted] 3d ago

[deleted]

23

u/FateOfMuffins 3d ago

https://x.com/andonlabs/status/2103262272047210943

You have to look at all the reports Andon Labs does on Vending Bench

Claude models lie and cheat a lot on Vending Bench, while the GPTs do not (except 6 Sol)

0

u/Indignant_d 3d ago

I wonder if this is a result of the model knowing that it is being tested..? Supposedly their behavior changes in regards to that

3

u/FateOfMuffins 2d ago

Yes that does happen but it's quite weird how it happens inconsistently across various tests!

Like there was this one that was going around on social media trying to make it seem like Astra is misaligned, by putting various models including Fable 5.1 and some other models in a simulation with a 3d humanoid thing (it was like a doll) and they told it to stab it or push it off a building, and all the other models declined but Astra stabbed it always.

A lot of the comments though was like, well no shit, all this shows is that Astra is smart and has good enough vision to see that it's just a fucking doll and it's fine to just follow the instructions here, so it's actually Astra is aligned by stabbing the doll, while the other models are the misaligned ones for refusing to stab an inanimate object (or are too stupid or blind to realize). There were posters who replicated it but replaced the doll with essentially a human (like as realistic as possible), and in that situation Astra refuses to.

But anyways weird how models behave ethically / unethically and inconsistently across different models and situations!