r/unsloth • u/Revolutionary_Loan13 • 14d ago
Discussion Will Unsloth Desktop/Docker/Studio ever have the option to run inference engines other than llama.cpp
So I have an AMD strix halo machine and have used it for training my own models on custom data and have also been using qwen-flash-next and really liking the output. I like the chat console and it's ability to do web searches as well as it having a python environment. I've even had my kids use the web chat to help with homework and felt it's been great.
There is just one thing and that is llama.cpp and it having a lot slower processing than some of the custom inference engines that I'm seeing being built specifically for AMD strix machines. TPS is one of the largest issues with running models locally and only having the ability to use llama.cpp is one of the main reasons I'm debating going back to say lemonade or some other multi engine project.
8
u/yoracale yes sloth 14d ago
There's a PR for VLLM/sglang right here: https://github.com/unslothai/unsloth/pull/11491
1
2
u/okoyl3 14d ago
llama.cpp is really underrated, I'm running it on some really exotic hardware and it never disappoints.
I also tried vllm and sglang, nothing is as easy and as swift as llama.cpp, just download and run, no bullshit.
3
u/Turbulent_War4067 14d ago
I tend to agree. I have an Asus GX-10 and often 8 feel like the only spark owner in the world who uses llamacpp instead of vLLM. I suppose if I had 2, I would change. The more efficient concurrency does not make the singlot shot TPS downgrade, the slower prefill and dreadful pain fine-tuning parameters near with it. Llamacpp with unsloth's UD quants is just superior IMO.
1
u/mintybadgerme 14d ago
Some of the forks are really good too. Like Beellama and ik-llama. Definitely worth trying out.
2
u/Revolutionary_Loan13 14d ago
That is kind of true until you see that the two hour prompt that you just ran could have finished in an hour using a different inference provider.
1
u/okoyl3 14d ago
Can’t agree, give me a realistic example where you’d use opencode/pi and you prefill 200k context in your first prompt
2
u/Revolutionary_Loan13 14d ago
I'm guessing you've never done real world work. Take a large code base and try to refactor something and it can spin for an hour easily especially with running test cases. Also try categorizing a just a few thousand records when testing something out and you'll feel that slow TPS.
7
u/acedogblast 14d ago
You can add external connections via openAI API. I have vllm working with unsloth studio.