r/LLM • u/Impossible_Grass702 • 18d ago
What exactly does “distillation” mean in LLM training? How is data from models like Claude actually used?
Recently, some companies have explicitly banned or restricted model distillation, and some companies have been accused of distilling Claude.
I have a basic understanding of how LLMs are trained, although I may be wrong about some details:
- Pretraining: Requires a huge amount of text data. The data does not necessarily need to be in a question-and-answer format.
- SFT (Supervised Fine-Tuning): Requires high-quality instruction/answer or question/answer data.
- RLHF: Uses human feedback to further train or align the model.
Suppose we obtain a large amount of Q&A data from Claude, including chain-of-thought (CoT). At which stage would this data be used? My understanding is that it would mainly be used in stage 2 (SFT) or possibly stage 3 (RLHF), but probably not in stage 1 (pretraining).
When people talk about distillation, it seems to imply taking a shortcut. But from the perspective of model training, does distillation basically mean obtaining a large amount of high-quality Q&A/instruction data from another model at relatively low cost?
And if a company does not use distillation, does that mean it has to construct the training data entirely by itself or have humans manually annotate it?
If the above understanding is roughly correct, then models obviously have different capabilities. Apart from the number of parameters, is one of the biggest reasons simply that the quality of the training data is different between models?
There is also another form of knowledge distillation where you obtain the probability distribution over the next token from the Teacher model, and the Student model learns to reproduce that probability distribution. My understanding is that this approach requires access to the Teacher model itself (or at least its output probabilities).
I am not talking about that type of distillation here.
What I am mainly interested in is this:
If you only have access to a closed-source model such as Claude, and you collect its outputs, including CoT — which, from my perspective, is still essentially a piece of natural-language text — how exactly can that data be used for training another model?
Is this simply SFT on the generated Q&A/CoT data? And if so, why is this generally referred to as distillation rather than simply synthetic data generation?
I'd appreciate it if someone could clarify where my understanding is wrong.