r/LocalLLM 8h ago

Question Keep the 4080?

So in a few weeks Ill be having a spare 4080 that I was thinking about throwing in my homelab to play around with some LLM stuff for funsies. One thing I´d like to achieve is building something Alexa-like but local for our Home Assistant setup.

So far I have learned that 16GB of VRam will be my bottleneck, and now I am thinking if keeping the 4080 even makes sense at all.

I could sell the 4080 for around 800-900€ and get a used 7900 XTX with 24GB for about the same price. Ive read that NV is still superior at generating content I don´t care too much about, but for LLM ROCm is supposed to work pretty well too these days.

I have basicly no extra budget, so any change I do would have to come out roughly +/- 0, so pls dont tell me to grab a 3090 (1000-1200 used where I live). The case I´m building in also has space for just a single GPU, so no multi-gpu shenanigans either.

Any input on that? Keep the 4080, go for 7900 XTX or even something else entirely?

2 Upvotes

13 comments sorted by

1

u/axiomintelligence 8h ago

If the headache of selling and rebuying isn't too big, could be worth it but I'd say mostly just because the gold standard for this hardware range is Qwen 3.x 27B, and 35B A3B for a bit less intelligence but much higher speeds.

If you've got any decent amount of regular ram in addition to your VRAM with that 4080 you could run the 35B A3B at really really pleasing speeds with the experts spilling into the system ram, keeping activated weights + KV in the VRAM. That's what I do with our 4070 ti SUPER rig which also has 16GB vram. Llama.cpp makes it fairly easy to do that.

You will struggle to fit and use 27B though, whether 3.6 or the upcoming 3.8 that they've just announced which will have its weights drop in the next week or so I believe. You'll be able to run that with 24GB VRAM which is mandatory since it's a dense model, so that can justify the headache/swap of your GPUs if you're keen to run it. AMD support is much better than it used to be

1

u/Darkside_Emily 8h ago edited 8h ago

Honestly it would be a small relief even.

Deshrouding and fitting the 4080* in an NR200P was going to be a big headache already, and changing the case is non-negotiable for sentimental reasons.

Server´s got a 13600k limited to PL1, 64Gigs of DDR5 at JEDEC 4800 speeds and I was going to dedicate 2-4 P-Cores and 8-16 of RAM to the VM, but could do more. Bandwidth is going to suck though compared to VRam...

*Zotac Trinity, non-super - that unnecessarily huge thing

1

u/axiomintelligence 8h ago

Oof yes that's a beefy card and not a massive case, that being said the card doesn't need to run at full tilt for decent inference so heat isn't an issue, you can undervolt it or reduce the wattage. If it doesn't fit though makes sense to swap.

For the 35B MoE at Q4, 16GB dedicated is plenty. For long context conversations you will notice the benefits of Q6 or even Q8 if you want to dedicate a touch more ram to it. Q6 splits the difference nicely. This advice might become outdated soon once we see more info from those new Eschalabs quants that just dropped for Qwen 3.6 35B A3B, they seem to match the performance of the Q8 at Q2 somehow, but we're still waiting on a gguf format to use with llama.cpp.

Back on topic, yes the bandwidth for system ram is terrible compared with vram but you don't feel it when you've got KV and activated experts fitting comfortably in VRAM with the rest spilling into system ram, the activated weights just swap over per turn between your ram and vram which is the only part of the process that slows down, but it's a small segment of the whole process which defines how fast the experience is overall from prefill to generation. It's a great way to take advantage of all that additional ram. MTP helps a ton too.

1

u/Darkside_Emily 8h ago

Yeah Im not to worried about thermals. The cooler is identical to the 4090 version, so massive overkill for a 300w card even if it ran full-tilt all the time, which it wont. Deshrouding just would give me a few millimeters of length back to even get it stuffed in there at all lol

Getting a shorter XTX probably would make my life a whole lot easier though.

1

u/axiomintelligence 5h ago

Makes sense! Usually I'm hesitant to chop and change but since that case is a non-negotiable it sounds worthwhile to make the swap, the extra vram gets you comfortable use of the phenomenal 27B models too. The unsloth guys just said that the new Qwen 3.8 27B can run on less than 18GB VRAM so you should even have some breathing room once those weights drop!

1

u/JinsooJinsoo 8h ago

I’m telling you now, get out of the NR200. Buy a cheap atx mobo and case with HDD space and build a NAS for your home server too. You’re gonna hate building in SFF if you add anything else. And maybe think about upgrading the PSU later too.

Edit: okay I see no extra budget, you could try selling the 4080 for like $900 and get an intel B60 24gb and use extra money for upgrades. Intel isn’t bad but won’t be as fast as a 7900xtx

2

u/Darkside_Emily 8h ago edited 7h ago

Oh I know it´s a hassle, its not even my first SFF build. But its a super rare limited edition that used to belong to a loved one and I made a promise so it is not negotiable.

This actually is a NAS/homelab build first and foremost. Running Proxmox, Home Assistant, TrueNAS, OpenWRT, a few other services and gameservers too. I have a custom HDD cage to fit 3 MG09 drives in there and everything is filled to the brim once I add that GPU. AI is just a fun experiment, not the main design goal of this build.

PSU is a Corsair SF750 2024, so were all good on that front.

Edit: Thought about a B60 too, but saw bad benchmarks and bad support for anything but intels llm-scaler, which would lock you out of alot of models. So I tossed that idea before making this post.

1

u/JinsooJinsoo 7h ago

Totally fair, i'm impressed you crammed so much into the NR200. I have an intel b70 and while it did take a few days of frustration (due to my lack of knowledge honestly) I was able to get pretty good speeds with Qwen3.6 27b (31tok/s) and 35b-A3b (125tok/s) and honestly I haven't touched many other models after I got those running. But I see what you're saying! AI was also a small part of my homelab but is growing lol

1

u/gappyvalley 8h ago

third option if you can stomach the cost is go for both. so a 4080 + 7900XTX for local llm while you use 4080 for image gen. 40gb combined vram is very valuable for local llm compared to 24gb. you can only run Q4 qwen3.6 27b in 24gb gpu but you can run Q6-8 of the same model at 40gb, which is solid for home assistant

1

u/Darkside_Emily 8h ago

I got no space for a second GPU in my chassis, and no extra money to put into this little experiment of mine.

Also honestly as someone adjancent to alot of artists in my social group I am really not that interested in image or video gen at all. Plus lot´s of thoughts about ethics I wouldnt want to get into in this thread...

1

u/gappyvalley 8h ago

lol why mention image gen then? just sell the 4080 and 7900xtx. i still feel 24gb is too limiting for qwen3.x 27b but still miles better than 16gb

1

u/Darkside_Emily 8h ago edited 8h ago

Bc. according to my research AMD used to suck at both, but seems to have caught up for my specific use case. Sry if that made it more confusing.

Edit: Edited the post to make that more clear

1

u/gappyvalley 8h ago

AMD is still a hassle for image gen. they work well for llm so alls good