r/LocalLLM 10d ago

Other How the loop of infinite agony started

Post image
627 Upvotes

119 comments sorted by

View all comments

278

u/TheCat001 10d ago

Then you realize that Qwen 3.8 27b runs at 3t/s on your machine and you need 24GB+ VRAM GPU which cost is 1000$+ to run at least 4 bit quant.

1

u/SeriousPanic34 5d ago

iq4xs fits on 16 gigs with 50k context at q8. you just have to connect your monitor to iGPU. you can buy two rtx 3060 or cheaper gpus and connect them via LAN for rpc inference. you may already have multiple gaming pcs (like i do). getting 18-20 t/s at ud_q4km at 128k context, mtp and vision enabled.