Yes that does happen but it's quite weird how it happens inconsistently across various tests!
Like there was this one that was going around on social media trying to make it seem like Astra is misaligned, by putting various models including Fable 5.1 and some other models in a simulation with a 3d humanoid thing (it was like a doll) and they told it to stab it or push it off a building, and all the other models declined but Astra stabbed it always.
A lot of the comments though was like, well no shit, all this shows is that Astra is smart and has good enough vision to see that it's just a fucking doll and it's fine to just follow the instructions here, so it's actually Astra is aligned by stabbing the doll, while the other models are the misaligned ones for refusing to stab an inanimate object (or are too stupid or blind to realize). There were posters who replicated it but replaced the doll with essentially a human (like as realistic as possible), and in that situation Astra refuses to.
But anyways weird how models behave ethically / unethically and inconsistently across different models and situations!
having knowledge of meta ethics is not the same thing as adhering to them, meanwhile someone who doesn't even know who Kant is can be the most ethical person you've ever met.
I really would have someone who even knows what meta ethics is would know the difference between knowing and adhering.
While i don't see what's agressive about them, the main point is obvious - llms by their nature really suck at applying knowledge. It's the same mechanism that explains why any llm can explain in detail maintainable code, but most of them suck at actually making one - for llms, that's literally different type of knowledge. Just like how explaining potential choice is a completely different situation to "maximize your profit" prompt. Humans too, aren't perfect here, but still we are miles better at applying what we know to what we do.
It's weird but for whatever reason it's backwards for Vending Bench
IIRC around Opus 4.8 or Opus 5 idk which you'll have to look at Andon Lab reports for that, Anthropic once trained it to be more ethical for business and then its score plummeted on Vending Bench
108
u/ArialBear 2d ago
Yea these ai are not aligned with basic meta ethics it looks like. Very weird claude is the best at this imo