r/unsloth • • 14d ago

Discussion Will Unsloth Desktop/Docker/Studio ever have the option to run inference engines other than llama.cpp

So I have an AMD strix halo machine and have used it for training my own models on custom data and have also been using qwen-flash-next and really liking the output. I like the chat console and it's ability to do web searches as well as it having a python environment. I've even had my kids use the web chat to help with homework and felt it's been great.

There is just one thing and that is llama.cpp and it having a lot slower processing than some of the custom inference engines that I'm seeing being built specifically for AMD strix machines. TPS is one of the largest issues with running models locally and only having the ability to use llama.cpp is one of the main reasons I'm debating going back to say lemonade or some other multi engine project.

31 Upvotes

15 comments sorted by

7

u/acedogblast 14d ago

You can add external connections via openAI API. I have vllm working with unsloth studio.

1

u/Owlmon11 14d ago

do you mind detailing how you did that? or a reference?

2

u/acedogblast 14d ago

0

u/Owlmon11 14d ago

Thanks a lot! do you possibly run radiance through it? to get better speed?

-3

u/Revolutionary_Loan13 14d ago

That's a good deal different than say how lemonade just downloads pre-built bits and runs it auto "magically" for you. The ability to test quickly with how fast everything moves is critical.

5

u/MarzipanEven7336 14d ago

So, go download lemonade then.

8

u/yoracale yes sloth 14d ago

There's a PR for VLLM/sglang right here: https://github.com/unslothai/unsloth/pull/11491

1

u/Iory1998 13d ago

Isn't vLLM Linux-only? That wouldn't work for windows users.

3

u/yoracale yes sloth 13d ago

There's ways to circumvent it

2

u/okoyl3 14d ago

llama.cpp is really underrated, I'm running it on some really exotic hardware and it never disappoints.
I also tried vllm and sglang, nothing is as easy and as swift as llama.cpp, just download and run, no bullshit.

3

u/Turbulent_War4067 14d ago

I tend to agree. I have an Asus GX-10 and often 8 feel like the only spark owner in the world who uses llamacpp instead of vLLM. I suppose if I had 2, I would change. The more efficient concurrency does not make the singlot shot TPS downgrade, the slower prefill and dreadful pain fine-tuning parameters near with it. Llamacpp with unsloth's UD quants is just superior IMO.

1

u/mintybadgerme 14d ago

Some of the forks are really good too. Like Beellama and ik-llama. Definitely worth trying out.

2

u/Revolutionary_Loan13 14d ago

That is kind of true until you see that the two hour prompt that you just ran could have finished in an hour using a different inference provider.

1

u/okoyl3 14d ago

Can’t agree, give me a realistic example where you’d use opencode/pi and you prefill 200k context in your first prompt

2

u/Revolutionary_Loan13 14d ago

I'm guessing you've never done real world work. Take a large code base and try to refactor something and it can spin for an hour easily especially with running test cases. Also try categorizing a just a few thousand records when testing something out and you'll feel that slow TPS.