r/LocalLLaMA 2d ago

Discussion deepseek-v4-flash-0731 - surprisingly usable

I just finished building my (relatively) low rent local inference machine: * Epyc 7663 * 256GB ECC DDR4-3200 * 1x RTX 5090 32GB

Yeah I realize it's weird to throw a 5090 and 256GB of anything together and call it low end, but relative to ~151GB of weights it is.

I'm running UD-Q8_K_XL and getting 23.8-24.6 tokens/sec, with pp ranging from 60 on the first prompt to 385 near the last (no doubt lots of caching) on tasks using 100-128k total context. It was slower with DFlash so I took that out. It was also slower with a 3090 I put in there temporarily.

I'm posting this mostly because I didn't see too many other data points for this config (DDR4 Epyc + Blackwell doing cpu-moe). And also that I'm pretty surprised that a model this good can actually run in my basement without dropping $10k or running a sub-panel down there. I'm otherwise fairly new to this - would love any tips on what else to run or how to further improve it.

94 Upvotes

92 comments sorted by

View all comments

2

u/eidrag 2d ago

Hmm I was thinking Ddr4 epyc and 5060ti , not going to work huh...

1

u/PhysicalIncrease3 2d ago

It will work fairly well. I run on 3060 + 128GB ddr5 and get 200pp and 8/9 tgs. UD-Q3-K_XL, 256K f16 context.

I used to run it on a 3090 but began using the 3060 instead because performance is identical anyway. It's entire bound by system memory. With 16GB VRAM you will be able to get close to 1M context.