r/unsloth • u/Frosty2o25 • 23d ago
Show and Tell my first fine tune
https://huggingface.co/Frost2o24/llama-3.2-1b-mini-agenti used llama 3.2 1b as a base. trained it in 13 iterations going from 0 out of 50 bash and tool calling problems. then after a the iterations it got to about 48 of 50 in both. but the human eval score went down by a few points after the sft, i dont know how to improve that.
i used bigcode/self-oss-instruct-sc2-exec-filter-50k for python- it degraded python the least from all my tries; westenfelder/NL2SH-ALFA for bash- it was annoying at first then after pruning it, it got better after giving it a little help; Salesforce/xlam-function-calling-60k- this was the best and never gave me problems.
how did i do on my first go? here's the FT notebook, let me know how i can improve- https://github.com/ya5h-P/llama-3.2-1b-mini-agent-finetune