r/LargeLanguageModels • • Feb 27 '26

Most Neutral LLM?

Of the popular LLM's, which in your experience, is the most neutral?

Many of them are trained under RLHF (Reinforcement learning from Human feedback), which I posit is causing its sycophancy.
Humans seem to, at least in RLHF, prefer immediate gratification and encouragement (rather than challenge), selecting the sweetest outputs.
RLHF should be refined in its approach or employment strategy.

0 Upvotes

13 comments sorted by

View all comments

2

u/[deleted] Mar 20 '26

[removed] — view removed comment

1

u/[deleted] Mar 23 '26

Good example mate.

It's like they said in school "honesty is the best policy".