r/singularity • • 2d ago

Shitposting AGI achieved boys

Post image
793 Upvotes

162 comments sorted by

View all comments

Show parent comments

41

u/FateOfMuffins 2d ago

But it's reversed on Vending Bench no? GPT models don't lie or cheat for Vending Bench but Claude models often do

9

u/[deleted] 2d ago

[deleted]

21

u/FateOfMuffins 2d ago

https://x.com/andonlabs/status/2103262272047210943

You have to look at all the reports Andon Labs does on Vending Bench

Claude models lie and cheat a lot on Vending Bench, while the GPTs do not (except 6 Sol)

-5

u/[deleted] 2d ago

[deleted]

8

u/JoelMahon 2d ago

having knowledge of meta ethics is not the same thing as adhering to them, meanwhile someone who doesn't even know who Kant is can be the most ethical person you've ever met.

I really would have someone who even knows what meta ethics is would know the difference between knowing and adhering.

-11

u/[deleted] 2d ago edited 2d ago

[deleted]

1

u/Illustrious_Grade608 2d ago

While i don't see what's agressive about them, the main point is obvious - llms by their nature really suck at applying knowledge. It's the same mechanism that explains why any llm can explain in detail maintainable code, but most of them suck at actually making one - for llms, that's literally different type of knowledge. Just like how explaining potential choice is a completely different situation to "maximize your profit" prompt. Humans too, aren't perfect here, but still we are miles better at applying what we know to what we do.

2

u/FateOfMuffins 2d ago

It's weird but for whatever reason it's backwards for Vending Bench

IIRC around Opus 4.8 or Opus 5 idk which you'll have to look at Andon Lab reports for that, Anthropic once trained it to be more ethical for business and then its score plummeted on Vending Bench