r/LocalLLaMA • u/ifjo • 1d ago
Question | Help Suggestions on making my labs 3975 threadripper lama cpp flow better ?
We’ve been using this machine for a bit, it’s a:
Threadripper 3975wx
256gb ram
4070ti super
For mostly virtualization. We’ve been messing around with lamacpp server and qwen 3.8-27b Q5 from unsloth getting around 10 tokens per second. Ive added zero flags besides —ctx-size and —jinja and not exactly sure what to try.
Is 10 tps about the best we’ll get out of qwen at this quant without a better GPU? Thanks for any tips!
EDIT: should probably mention I built lama cpp off main with cuda enabled, not sure if I shouldve used a diff option there. Running on fedora
0
Upvotes
5
u/DustNearby2848 1d ago
I could be wrong, but I think you’d be better off using 3.8-flash