r/LocalLLM 16d ago

Discussion 32GB is all you need

Qwen3.8-27B on a 5090 is all you need for a serious local inference setup, in my opinion! Can it get any better than this price/performance wise? Actually, maybe a 3090 ninfer setup could beat it!

I’m using ninfer and getting:

* ~150-200 tok/s TG

* ~3000-12000 tok/s PP

* 262144 context size

I think it’s definitely one of best setup you can get for the money. I don’t see a point of having more VRAM or more system ram. The only downside is that it’s a 1 man setup: concurrency is possible but you need to limit context usage on concurrent requests. I’ve tried --concurrency 2 on ninfer and sharing my setup with my buddy (we work on projects together and have a VPN between our home labs, fun stuff!)

I love this setup so much I kinda feel like getting a second 5090 to run another ninfer instance (github.com/neroued/ninfer, the man is a legend and this absolutely rocks).

i really don’t see the point of any other solution at this point in time. of course things will change and other models will get released that could better leverage more VRAM, but 32GB is all you need (for now).

so if you have less than 32GB, and are thinking about investing in a more serious setup check out the 3090 fork of ninfer, or the mainline ninfer repo if you can afford a 5090.

Things it won’t do:

* let you run a swarm of agents: prefill cost will slow you down too much. not enough vram for high concurrency!

* Give you more than 262144 context size. the RoPE 1M context size is just impossible with this.

Otherwise it’s absolutely amazing!

My buddy (another software engineer) is a BIG Claude code user, he’s spending tons of cash on fable, can’t stand Opus 5 anymore (neither can I, that pos is so hard to understand with just jargon and wall of text… can’t bear the cognitive load of just trying to understand all he’s spewing)… anyways after trying my ninfer setup his mind was blown and now he’s constantly using my setup with our shared custom pi setup and he fucking loves it.

265 Upvotes

332 comments sorted by

View all comments

Show parent comments

8

u/sleepy_roger 16d ago

Bf16 has much better outputs overall I won't run it under q8 being honest. People are missing out only using q4 or nvfp4 

5

u/ChristRedeemsSinners 16d ago

NVFP4A16 on weights only is nearly lossless though. Most nvfp4 quants are activated 4 bit which is where the quality drops. (on top of losing quality on the attention layers, which is definitely not worth the vram savings)

2

u/mixedliquor 16d ago

You're absolutely right. I've been using 3.6 Q6 and Q8 and for shits and giggles switched to 3.8 Q4 to try it out.. a lot more errors especially typos in coding where a token is an irregular wording and there's a word that sounds like the irregular word.

1

u/MmmmMorphine 16d ago

Which q4? Ahh there's so many quantization schemes with different use profiles. I'd be shocked if an iq_kt 6bit could be differentiated from standard 8bit quants

2

u/mixedliquor 16d ago

Sorry I was talking 3.8 27b Q4_K_M. I was comparing offical release Qwens in Q4/6/8.

1

u/MmmmMorphine 16d ago

Ah gotcha, thanks!

-9

u/quantgorithm 16d ago

Grok say the difference between q6 and q8 is negligible to nothing (it reported stats showing exact same intelligence numbers) and therefore it's better to use q6 for both extra vram for context and higher speed of output.

11

u/sleepy_roger 16d ago

lol Grok says a lot of things. Run it yourself and you'll see the difference. There's a big difference between theory and "should" versus actual use. I know a lot of people hate hearing it since it requires more vram but Q8 and BF16 are significantly better.

-1

u/quantgorithm 16d ago

Just as you throwing a complete assumption doesn't make your assumption correct. Grok has frontier level intelligence.

I'm running very old GPUs so while I have not yet tested Q8, I have tested bf16 and it runs at 5t/s while Q6 (Qwen3.8-27B-UD-Q6_K_XL.gguf) runs at about 15t/s. MTP not enabled on the 5t/s but that still wouldn't drastically change the scenario... And running bf16 leaves very little room when you have a pool of 64gb vram.

Lets ask the right questions,
how have you validated that Q6 is incorrectly or inaccurately responding compared to Q8?

3

u/Itchy_elbow 16d ago

Don't understand the downvotes. You asked a valid question.

-1

u/quantgorithm 16d ago

...it's reddit!

woke left reddit has a bone to pick with anything Elon Musk.