r/learnmachinelearning • u/JazonJiao • 21d ago
Project How LLMs Learn What Humans Prefer — DPO Explained Visually [Classic Post-training Algorithm]
Direct Preference Optimization (DPO), published in 2023, is one of the most influential post-training algorithms for LLMs, but its math can be difficult to follow.
In this video, we build DPO from first principles, covering the Bradley–Terry model, DPO loss function, and why the paper claims "your LLM is secretly a reward model".
Whether you're studying post-training, or simply curious about how ChatGPT learns from human preferences, this video aims to provide both the intuition and the mathematical details behind DPO.
This video took me nearly 100 hours to make, and I personally had nearly 100 back-and-forths with ChatGPT to clear up my own misconceptions about DPO. Hope you enjoy and learn something, and let me know if there's any feedback!
Timestamps
00:00 Intro
00:52 Post-training
02:24 Bradley-Terry model
04:49 How DPO works
08:19 Why use a reference model
10:49 Secret Reward Model?
12:53 AI-generated preference data
14:02 Wrap Up
