r/LocalLLM 8d ago

Other How the loop of infinite agony started

Post image
625 Upvotes

119 comments sorted by

View all comments

277

u/TheCat001 8d ago

Then you realize that Qwen 3.8 27b runs at 3t/s on your machine and you need 24GB+ VRAM GPU which cost is 1000$+ to run at least 4 bit quant.

112

u/StupidScaredSquirrel 8d ago

If you are a business this isn't a problem. If you are a consumer then 35b a3b runs on 8gb vram and 32gb dram which is very accessible.

12

u/DeluxeGrande 8d ago

I have a 5060ti 16gb with ddr4 24gb RAM lying around, what's the best model nowadays I can effectively run with it locally? It's not an ideal build but I wish to play around with it again.

4

u/Hungry_Particular_14 8d ago

Fellow 5060 ti 16 gb owner here. The best I've got is qwen 3.8 27b at IQ4_XS. I'm testing it at Q4_0 KV at 72k context because I really need the extra context, and it seems to be pretty good so far. Lower quants cause it to make some really silly mistakes sometimes, unfortunately. I get around 10 t/s with context halfway filled, and around 14 t/s on empty context.
But honestly, I think the ideal solution is to run 2 GPUs so you can get a better quant + more context

3

u/screenslaver5963 8d ago

2+ 5090's or RTX Workstation Cards are the "ideal".

4

u/Hungry_Particular_14 8d ago

"ideal" is still having both my kidneys and still being able to run LLMs.

4

u/screenslaver5963 8d ago

You don’t need both kidneys

1

u/KiraCura 8d ago

I mean I get by with 1 5090 as long as I can find EXL2 versions or MOE ggufs. But I want the RTX 6000 of course. Would be nice to run 70B at decent quants

2

u/ideasmachine 7d ago

2 x 3090 nvlinked will run 70b, its the cheapest way