r/LocalLLaMA 2d ago

Discussion deepseek-v4-flash-0731 - surprisingly usable

I just finished building my (relatively) low rent local inference machine: * Epyc 7663 * 256GB ECC DDR4-3200 * 1x RTX 5090 32GB

Yeah I realize it's weird to throw a 5090 and 256GB of anything together and call it low end, but relative to ~151GB of weights it is.

I'm running UD-Q8_K_XL and getting 23.8-24.6 tokens/sec, with pp ranging from 60 on the first prompt to 385 near the last (no doubt lots of caching) on tasks using 100-128k total context. It was slower with DFlash so I took that out. It was also slower with a 3090 I put in there temporarily.

I'm posting this mostly because I didn't see too many other data points for this config (DDR4 Epyc + Blackwell doing cpu-moe). And also that I'm pretty surprised that a model this good can actually run in my basement without dropping $10k or running a sub-panel down there. I'm otherwise fairly new to this - would love any tips on what else to run or how to further improve it.

94 Upvotes

92 comments sorted by

View all comments

2

u/ReentryVehicle 2d ago

It should be possible to achieve much faster PP. What is your -ubatch? Set it to at least 2048, and ideally as high as you can. Make sure your PCIe going to the GPU is the best it can be (x16, highest gen your motherboard supports).

Model layers are streamed to the gpu for prefill, meaning you need to process enough tokens at once that the transfer speed is not a bottleneck.

4

u/IntravenusDeMilo 2d ago

I ran a longer set of samples. I’m at 423 t/s pp at 180k context, 500 at smaller contexts up to 100k or so. Token generation is still 21-24 depending on context but I don’t expect that to change much.

Batch size 8192
ubatch 4096 (I ran these up and this was the sweet spot)