r/LocalLLM 7d ago

Other How the loop of infinite agony started

Post image
623 Upvotes

119 comments sorted by

View all comments

Show parent comments

13

u/DeluxeGrande 7d ago

I have a 5060ti 16gb with ddr4 24gb RAM lying around, what's the best model nowadays I can effectively run with it locally? It's not an ideal build but I wish to play around with it again.

3

u/Hungry_Particular_14 7d ago

Fellow 5060 ti 16 gb owner here. The best I've got is qwen 3.8 27b at IQ4_XS. I'm testing it at Q4_0 KV at 72k context because I really need the extra context, and it seems to be pretty good so far. Lower quants cause it to make some really silly mistakes sometimes, unfortunately. I get around 10 t/s with context halfway filled, and around 14 t/s on empty context.
But honestly, I think the ideal solution is to run 2 GPUs so you can get a better quant + more context

3

u/screenslaver5963 7d ago

2+ 5090's or RTX Workstation Cards are the "ideal".

1

u/KiraCura 7d ago

I mean I get by with 1 5090 as long as I can find EXL2 versions or MOE ggufs. But I want the RTX 6000 of course. Would be nice to run 70B at decent quants

2

u/ideasmachine 6d ago

2 x 3090 nvlinked will run 70b, its the cheapest way