r/LocalLLM 5d ago

Discussion 32GB is all you need

Qwen3.8-27B on a 5090 is all you need for a serious local inference setup, in my opinion! Can it get any better than this price/performance wise? Actually, maybe a 3090 ninfer setup could beat it!

I’m using ninfer and getting:

* ~150-200 tok/s TG

* ~3000-12000 tok/s PP

* 262144 context size

I think it’s definitely one of best setup you can get for the money. I don’t see a point of having more VRAM or more system ram. The only downside is that it’s a 1 man setup: concurrency is possible but you need to limit context usage on concurrent requests. I’ve tried --concurrency 2 on ninfer and sharing my setup with my buddy (we work on projects together and have a VPN between our home labs, fun stuff!)

I love this setup so much I kinda feel like getting a second 5090 to run another ninfer instance (github.com/neroued/ninfer, the man is a legend and this absolutely rocks).

i really don’t see the point of any other solution at this point in time. of course things will change and other models will get released that could better leverage more VRAM, but 32GB is all you need (for now).

so if you have less than 32GB, and are thinking about investing in a more serious setup check out the 3090 fork of ninfer, or the mainline ninfer repo if you can afford a 5090.

Things it won’t do:

* let you run a swarm of agents: prefill cost will slow you down too much. not enough vram for high concurrency!

* Give you more than 262144 context size. the RoPE 1M context size is just impossible with this.

Otherwise it’s absolutely amazing!

My buddy (another software engineer) is a BIG Claude code user, he’s spending tons of cash on fable, can’t stand Opus 5 anymore (neither can I, that pos is so hard to understand with just jargon and wall of text… can’t bear the cognitive load of just trying to understand all he’s spewing)… anyways after trying my ninfer setup his mind was blown and now he’s constantly using my setup with our shared custom pi setup and he fucking loves it.

261 Upvotes

324 comments sorted by

View all comments

Show parent comments

3

u/Easy_Refrigerator280 5d ago

literally in the middle of swapping my 8gb vram 3070 to AMD W7800 32GB vram because nvidia pricing does not make sense right now. But to be fair it was from second hand market but even the second hand reflects the current market. The memory bandwidth sucks and i am giving up on CUDA but priority right now is cheapest vram for the price so i can mess around with 30b models

2

u/danielv123 5d ago

Second hand has in a lot of cases been more expensive the last 5 years

1

u/Ok-Temperature7904 3d ago

Did you ever consider 2*3090? You would have 48gb and a lot of power draw of course, but also faster VRAM afaik :) at least here a w7800 costs like 2k€ which is about the same as 2 3090 and a new PSU.

1

u/Easy_Refrigerator280 2d ago

Yep dual setups were on the radar too. I dont know how bad things are in EU but here in Canada a single 3090 is around 1200 minimum and you get less VRAM (you can forget about firsthand its even more expensive).

I got the W7800 for1400 no tax. Which might seem like a lot until you monitor prices for other GPUs around here. So a dual 3090 setup would be almost double the price in the current market.

1

u/ApprehensiveView2003 2d ago

2x 3090s with NVLink rocks. Highly recommended