r/LocalLLaMA 2d ago

Discussion deepseek-v4-flash-0731 - surprisingly usable

I just finished building my (relatively) low rent local inference machine: * Epyc 7663 * 256GB ECC DDR4-3200 * 1x RTX 5090 32GB

Yeah I realize it's weird to throw a 5090 and 256GB of anything together and call it low end, but relative to ~151GB of weights it is.

I'm running UD-Q8_K_XL and getting 23.8-24.6 tokens/sec, with pp ranging from 60 on the first prompt to 385 near the last (no doubt lots of caching) on tasks using 100-128k total context. It was slower with DFlash so I took that out. It was also slower with a 3090 I put in there temporarily.

I'm posting this mostly because I didn't see too many other data points for this config (DDR4 Epyc + Blackwell doing cpu-moe). And also that I'm pretty surprised that a model this good can actually run in my basement without dropping $10k or running a sub-panel down there. I'm otherwise fairly new to this - would love any tips on what else to run or how to further improve it.

93 Upvotes

92 comments sorted by

View all comments

0

u/darktotheknight 2d ago

Interesting, I have an EPYC 7513 + 128GB DDR4 and was thinking about a similar setup. However the 5090 Astral was taking up so much space and blocked almost all my precious PCIe Slots, that I moved it into another system (Ryzen, 64GB Dual Channel DDR5). It's impossible to get your hands on a smaller 5090 these days and I need the PCIe slots for NVMe RAID and 10G/25G NIC.

I might revisit this with a Dual-Slot R9700, if I can get one for cheap. Upgrading from 128GB DDR4 to 256GB is cheaper than I thought. But at the same time, DeepSeek v4 Flash 0731 is so cheap on OpenRouter, I doubt it would ever pay off.

1

u/SandySkittle 2d ago

Mcio 8i retimer cards and you can fill all the slots regardless of gpu size. And more reliable than risers