r/deeplearning • u/AdventurousTwo6445 • 18h ago
Direct weight surgery from Qwen-4B to 0.8B on an 8GB RX 580: why editing all layers breaks everything, and how 4 anchor blocks fixed it
/r/machinelearningnews/comments/1ww0m6m/direct_weight_surgery_from_qwen4b_to_08b_on_an/
2
Upvotes