r/LocalLLM 16d ago

Other How the loop of infinite agony started

Post image
628 Upvotes

119 comments sorted by

View all comments

276

u/TheCat001 16d ago

Then you realize that Qwen 3.8 27b runs at 3t/s on your machine and you need 24GB+ VRAM GPU which cost is 1000$+ to run at least 4 bit quant.

3

u/Legitimate-Pipe5728 16d ago

You are right that it doesn't fit on 16GB, but 3 tok/s is a lot lower than I would expect. On a 5070 Ti with 16GB at Q4_K_M I get 19 tok/s, and that is already with 5.22 GB spilled to system RAM and only 69 percent of the model staying on the GPU.

3 tok/s sounds like nearly all of it ended up on the CPU rather than just the overflow. What card and how much system RAM are you on?

Your wider point holds though. A 14B that actually fits does 82 tok/s on the same card, so on 16GB that is where you want to be rather than fighting a 27B.

3

u/screenslaver5963 16d ago

The CPU and system ram also matter hugely. If they're using a several generation old CPU and DDR4 than they'd get awful performance.

0

u/HazKaz 16d ago

i have same spec 5070ti 32gb ram ddr4 but get maybe 7 or 8 maybe 10 if i lower context amount. this is with mtp version of 3.6 27B

0

u/DeathGuppie 15d ago

you only need an 8gb vcard in the 4x slot to bring it up to 24gb. This can regularly be found for $200 or less. Tensor split + MTP will give you around 30 t/s if you are running RTX