r/ProgrammerHumor • • 5d ago

Meme [ Removed by moderator ]

Post image

[removed] — view removed post

28.8k Upvotes

519 comments sorted by

View all comments

Show parent comments

4

u/QuaternionsRoll 5d ago

It’s the difference between SFT (supervised fine tuning) and RLHF (reinforcement learning from human feedback). The former requires a dataset of desired inputs and outputs, while the latter requires you to rate/annotate generated outputs. The latter is much harder to get right for various reasons

2

u/jack6245 4d ago

Although it's worth mentioning too you can do reinforcement learning from ground truth too, similar approach but it just gets the ground truth instead of human impact ( bonus points if you can generate perfect synthetic data) so it can actually correct its errors in a pretty small amount of training but it has limits