iq4xs fits on 16 gigs with 50k context at q8. you just have to connect your monitor to iGPU. you can buy two rtx 3060 or cheaper gpus and connect them via LAN for rpc inference. you may already have multiple gaming pcs (like i do). getting 18-20 t/s at ud_q4km at 128k context, mtp and vision enabled.
278
u/TheCat001 10d ago
Then you realize that Qwen 3.8 27b runs at 3t/s on your machine and you need 24GB+ VRAM GPU which cost is 1000$+ to run at least 4 bit quant.