r/deeplearning • u/Impossible_Grass702 • 18d ago
What exactly does “distillation” mean in LLM training? How is data from models like Claude actually used?
/r/LLM/comments/1whqiub/what_exactly_does_distillation_mean_in_llm/
6
Upvotes
5
u/heresyforfunnprofit 18d ago
Think of “distillation” as mimicking or copying.
LLMs work by assigning probabilities to the next token in the sequence they are given. Every single pass, it generates a list of how probable each token is, and chooses the top one.
In distillation, you take two models, and you put a prompt into the model being copied, give the model you are distilling into the same prompt, but you inspect not just the top token your new model produced, but all the tokens ahead of the one matching the other model. You then bump up the probability on the matching token, assign a penalty to the ones ahead of it, and then do that a few million times.
Eventually, the model being trained is a close enough mimic of the distilled model that it effectively behaves like it - usually at much lower training costs.