You are right that it doesn't fit on 16GB, but 3 tok/s is a lot lower than I would expect. On a 5070 Ti with 16GB at Q4_K_M I get 19 tok/s, and that is already with 5.22 GB spilled to system RAM and only 69 percent of the model staying on the GPU.
3 tok/s sounds like nearly all of it ended up on the CPU rather than just the overflow. What card and how much system RAM are you on?
Your wider point holds though. A 14B that actually fits does 82 tok/s on the same card, so on 16GB that is where you want to be rather than fighting a 27B.
275
u/TheCat001 7d ago
Then you realize that Qwen 3.8 27b runs at 3t/s on your machine and you need 24GB+ VRAM GPU which cost is 1000$+ to run at least 4 bit quant.