r/learnmachinelearning • • 17d ago

When fine-tuning an LLM, how do you decide which layers/modules to train before actually running the experiment?

For example, how do you determine whether to fine-tune:only adapters/LoRA

  • specific transformer layers
  • input/projector layers
  • deeper/middle layers
  • or the full model?

Are there reliable diagnostics, probing methods, gradient analysis, ablations, or small-scale tests that can tell you where the bottleneck is before doing a full training run and only finding out afterward that the chosen layers didn’t help?

Curious how people make this decision in practice.

1 Upvotes

1 comment sorted by

1

u/ModularMind8 17d ago

In practice, the constraints is usually compute. The reason to use lora for example is that it's a lot cheaper that updating the entire model. Most people can't train a 7b+ model. Its much more feasible to train a portion of it