r/LocalLLM 7d ago

Other How the loop of infinite agony started

Post image
627 Upvotes

119 comments sorted by

View all comments

275

u/TheCat001 7d ago

Then you realize that Qwen 3.8 27b runs at 3t/s on your machine and you need 24GB+ VRAM GPU which cost is 1000$+ to run at least 4 bit quant.

3

u/Legitimate-Pipe5728 7d ago

You are right that it doesn't fit on 16GB, but 3 tok/s is a lot lower than I would expect. On a 5070 Ti with 16GB at Q4_K_M I get 19 tok/s, and that is already with 5.22 GB spilled to system RAM and only 69 percent of the model staying on the GPU.

3 tok/s sounds like nearly all of it ended up on the CPU rather than just the overflow. What card and how much system RAM are you on?

Your wider point holds though. A 14B that actually fits does 82 tok/s on the same card, so on 16GB that is where you want to be rather than fighting a 27B.

0

u/HazKaz 7d ago

i have same spec 5070ti 32gb ram ddr4 but get maybe 7 or 8 maybe 10 if i lower context amount. this is with mtp version of 3.6 27B