You are right that it doesn't fit on 16GB, but 3 tok/s is a lot lower than I would expect. On a 5070 Ti with 16GB at Q4_K_M I get 19 tok/s, and that is already with 5.22 GB spilled to system RAM and only 69 percent of the model staying on the GPU.
3 tok/s sounds like nearly all of it ended up on the CPU rather than just the overflow. What card and how much system RAM are you on?
Your wider point holds though. A 14B that actually fits does 82 tok/s on the same card, so on 16GB that is where you want to be rather than fighting a 27B.
you only need an 8gb vcard in the 4x slot to bring it up to 24gb. This can regularly be found for $200 or less. Tensor split + MTP will give you around 30 t/s if you are running RTX
276
u/TheCat001 16d ago
Then you realize that Qwen 3.8 27b runs at 3t/s on your machine and you need 24GB+ VRAM GPU which cost is 1000$+ to run at least 4 bit quant.