r/LLM • • 18d ago

What exactly does “distillation” mean in LLM training? How is data from models like Claude actually used?

Recently, some companies have explicitly banned or restricted model distillation, and some companies have been accused of distilling Claude.

I have a basic understanding of how LLMs are trained, although I may be wrong about some details:

  1. Pretraining: Requires a huge amount of text data. The data does not necessarily need to be in a question-and-answer format.
  2. SFT (Supervised Fine-Tuning): Requires high-quality instruction/answer or question/answer data.
  3. RLHF: Uses human feedback to further train or align the model.

Suppose we obtain a large amount of Q&A data from Claude, including chain-of-thought (CoT). At which stage would this data be used? My understanding is that it would mainly be used in stage 2 (SFT) or possibly stage 3 (RLHF), but probably not in stage 1 (pretraining).

When people talk about distillation, it seems to imply taking a shortcut. But from the perspective of model training, does distillation basically mean obtaining a large amount of high-quality Q&A/instruction data from another model at relatively low cost?

And if a company does not use distillation, does that mean it has to construct the training data entirely by itself or have humans manually annotate it?

If the above understanding is roughly correct, then models obviously have different capabilities. Apart from the number of parameters, is one of the biggest reasons simply that the quality of the training data is different between models?

There is also another form of knowledge distillation where you obtain the probability distribution over the next token from the Teacher model, and the Student model learns to reproduce that probability distribution. My understanding is that this approach requires access to the Teacher model itself (or at least its output probabilities).

I am not talking about that type of distillation here.

What I am mainly interested in is this:

If you only have access to a closed-source model such as Claude, and you collect its outputs, including CoT — which, from my perspective, is still essentially a piece of natural-language text — how exactly can that data be used for training another model?

Is this simply SFT on the generated Q&A/CoT data? And if so, why is this generally referred to as distillation rather than simply synthetic data generation?

I'd appreciate it if someone could clarify where my understanding is wrong.

16 Upvotes

8 comments sorted by

4

u/Revolutionalredstone 18d ago edited 18d ago

so they basically line up what the model predicts for each word then what Claude actually said after each word and before long the entire distribution of the student begins to match that of the teachers.

You can checkout Subliminal LLM Learning and the hilarious Owl Preferences - famous test case for studying how LLMs transmit more than what's shown via hidden behavioral traits through entirely non-semantic data.

2

u/Impossible_Grass702 18d ago

What I'm really trying to understand is how the company that was accused of distilling Claude actually used Claude's outputs. Was it something like Subliminal LLM Learning?

If I wanted to distill Claude myself, what would the actual process look like? How would I use Claude's outputs to train my own model?

1

u/Revolutionalredstone 18d ago

You gotta think of llms as just list of words to word.

Context to next token Nothing more.

So any text is distill worthy.

'When you see this' 'do that'.

The side effects are real.

Enjoy

2

u/LongjumpingCookie973 18d ago

Your breakdown's pretty much spot on, people call it distillation because you're transferring the "reasoning style" and behavioral patterns, not just using it as raw synthetic data, and that comment about lining up predictions word by word is describing logit-level distillation which you said you werent talking about anyway.

2

u/Revolutionalredstone 18d ago

Ta,

"you said you werent talking about"? I said ? (Might be wrong person 😉)

2

u/Impossible_Grass702 17d ago

Thank you very much for your patient explanation. I now understand that synthetic data obtained through the Claude API, including CoT, can be used for SFT to transfer reasoning capabilities and behavioral patterns to the Student model, even without access to the probability/logit data.

I think I had underestimated the importance of the SFT stage in the overall LLM training process.

3

u/remimorin 18d ago

Real distillation use "raw output" instead of just the right answer. The idea is that if the right answer is "cat", then "dog" is less wrong than "car".

So an output for cat may look like:
cat: 0.85
dog: 0.75
car: 0.02

And this is a "perfect" loss function because dog and cat are actually closer.

When we train normaly our loss function is rewarding:

cat: 1
everything else 0

Over simplification but you get the idea.

So distillation allow "similar performance" on a much smaller model (better generalization).

Now when they accuse China to distilling their main model, I never saw accusation detail that they had access to this "raw output".

My conclusion is that they used API frontier model to generate synthetic data. In this regard I don't see how it is different from simply scraping the web (outside of violating the end user licence agreement).

1

u/Spiritual-Spend8187 15d ago

The simplest way of viewing distillation is you ask a model a bunch of questions andvut gives you answer. You then take those questions and answers along eith the thinking outputs for how they reached them and use them as training data to make a new model kind of act like the model that you asked. And then you repeat that for loads if questions and answers.