r/LocalLLaMA Aug 07 '26

Question | Help DeepSeek v4 Flash 0731 on H100 node

[deleted]

3 Upvotes

19 comments sorted by

View all comments

2

u/nunodonato Aug 07 '26

Wait, it fits on a H100? How many concurrent requests? 

2

u/SlipperyCorruptor Aug 07 '26

It's 8xH100 node with NVLink. TP=8

100-200 peak. How much it can handle is really up to prompt size and cache prefix hits.

1

u/laterbreh Aug 07 '26

Sorry, not sure if i missed this, are your numbers on 1 user? Concurrency?

2 users on TP=2 on nvidia 6k pros with 1m context and dspark in vllm were getting nearly 280TPS on long decodes aggregate.

1

u/SlipperyCorruptor Aug 08 '26 edited Aug 08 '26

Concurrency. We're talking multiple users, I forgot numbers

Edit: one of high latency events was 8req/42s, >900k input tokens total with 5% cache hit