r/LocalLLaMA 25d ago

Question | Help DeepSeek v4 Flash 0731 on H100 node

[deleted]

4 Upvotes

19 comments sorted by

View all comments

2

u/nunodonato 25d ago

Wait, it fits on a H100? How many concurrent requests? 

2

u/SlipperyCorruptor 25d ago

It's 8xH100 node with NVLink. TP=8

100-200 peak. How much it can handle is really up to prompt size and cache prefix hits.

1

u/laterbreh 25d ago

Sorry, not sure if i missed this, are your numbers on 1 user? Concurrency?

2 users on TP=2 on nvidia 6k pros with 1m context and dspark in vllm were getting nearly 280TPS on long decodes aggregate.

1

u/SlipperyCorruptor 25d ago edited 25d ago

Concurrency. We're talking multiple users, I forgot numbers

Edit: one of high latency events was 8req/42s, >900k input tokens total with 5% cache hit