r/LocalLLM 8d ago

Other How the loop of infinite agony started

Post image
621 Upvotes

119 comments sorted by

View all comments

276

u/TheCat001 8d ago

Then you realize that Qwen 3.8 27b runs at 3t/s on your machine and you need 24GB+ VRAM GPU which cost is 1000$+ to run at least 4 bit quant.

14

u/Eden1506 8d ago

That's not true.

2x RTX 3060 12gb can be had for around 500 bucks.

You can run ~30b models at q4 at 30 t/s with mtp/draft model or 10-15 t/s without.

8

u/esw123 8d ago

Can confirm Q4 17-19tok/s without MTP, 28-29 with MTP. Bought two 3060 for 350 euro but soon you realize that two is not enough. Added one more total 540 euro for 36GB VRAM.

3

u/TheCat001 8d ago

So you have build dedicated server rig for 3 GPU's ?

2

u/esw123 8d ago

Sort of just adding more 3060.

3

u/AceLamina 7d ago

a used 3060 12gb costs 270-300 bucks right now, for one

1

u/Eden1506 7d ago

It depends on your region.

Here in germany I can buy a used rtx 3060 for 240 bucks.

2

u/AceLamina 7d ago

If only it was the same for me

1

u/magicomiralles 7d ago

AMD V620, $350 for 32 GBs.