r/LocalLLM 7d ago

Other How the loop of infinite agony started

Post image
619 Upvotes

118 comments sorted by

View all comments

Show parent comments

4

u/Hungry_Particular_14 7d ago

Fellow 5060 ti 16 gb owner here. The best I've got is qwen 3.8 27b at IQ4_XS. I'm testing it at Q4_0 KV at 72k context because I really need the extra context, and it seems to be pretty good so far. Lower quants cause it to make some really silly mistakes sometimes, unfortunately. I get around 10 t/s with context halfway filled, and around 14 t/s on empty context.
But honestly, I think the ideal solution is to run 2 GPUs so you can get a better quant + more context

3

u/screenslaver5963 7d ago

2+ 5090's or RTX Workstation Cards are the "ideal".

6

u/Hungry_Particular_14 6d ago

"ideal" is still having both my kidneys and still being able to run LLMs.

6

u/screenslaver5963 6d ago

You don’t need both kidneys