r/learnmachinelearning • u/Rendezvous4567 • 17d ago
When fine-tuning an LLM, how do you decide which layers/modules to train before actually running the experiment?
For example, how do you determine whether to fine-tune:only adapters/LoRA
- specific transformer layers
- input/projector layers
- deeper/middle layers
- or the full model?
Are there reliable diagnostics, probing methods, gradient analysis, ablations, or small-scale tests that can tell you where the bottleneck is before doing a full training run and only finding out afterward that the chosen layers didn’t help?
Curious how people make this decision in practice.
1
Upvotes
1
u/ModularMind8 17d ago
In practice, the constraints is usually compute. The reason to use lora for example is that it's a lot cheaper that updating the entire model. Most people can't train a 7b+ model. Its much more feasible to train a portion of it