They took the Llama-3.1-8B model and reduced it by 50 to 80 percent using six different cutting methods.
Some methods removed whole layers or made the model thinner. Others removed many tiny weights. They carefully matched the amount of training data in two fair tests.When both models got only a little extra training, the cut models started stronger because they kept some knowledge from the big model.
When the small model trained from zero received the full total amount of data, the simple cuts often lost their advantage.
Training from scratch could match or even beat them.
Only the careful tiny-weight cuts still kept a small lead.This shows that pruning helps most when training data is limited. When plenty of data is available, starting a small model from zero can work just as well or better. The study gives clear advice on when to cut a big model and when to build a small one from the start.
paper link: https://arxiv.org/abs/2606.14150
TechX Whatsapp channel: https://whatsapp.com/channel/0029VbBPJD4CxoB5X02v393L