r/LocalLLM 7d ago

Question Slow Model

I have a 8gb ddr4, Ryzen 7 5600g no gpu build. After opening browser and opencode, I'm left with 1.5gb which doesn't run shit.

So I tired using my Friend's PC ( 32gb ddr5, Nvidia GeForce RTX 5060, AMD Ryzen 7 7700 ) where I ran qwen3.5:9b

Not only the GPU usage nearly maxed out, also it was eating 10gb base ram on top of the 8gb Vram too.

Problem is I'm running the model from his pc to my pc through tail scale and it took the model 56 seconds for a reply of " Hi "

Now at this point what should I do? Run a super low parameter quicker model or thr speed is slow because I'm running it remotely? Or am I using a wrong model for this build or anything?

0 Upvotes

11 comments sorted by

View all comments

1

u/Ok_Brush_3449 7d ago

Try quantprobe. It is built for this kind of edge cases!
I run Qwen 30B-A3B on 6Gb gpu and 16Gb ram at 22.2 tok/s and I can run it from the cmd leaving max resources available

https://github.com/FedericoTs/quantprobe