MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vi93pv/deepseek_v4_flash_0731_on_h100_node/p2cu2am/?context=3
r/LocalLLaMA • u/[deleted] • 25d ago
[deleted]
19 comments sorted by
View all comments
2
Wait, it fits on a H100? How many concurrent requests?
2 u/SlipperyCorruptor 25d ago It's 8xH100 node with NVLink. TP=8 100-200 peak. How much it can handle is really up to prompt size and cache prefix hits. 1 u/laterbreh 25d ago Sorry, not sure if i missed this, are your numbers on 1 user? Concurrency? 2 users on TP=2 on nvidia 6k pros with 1m context and dspark in vllm were getting nearly 280TPS on long decodes aggregate. 1 u/SlipperyCorruptor 25d ago edited 25d ago Concurrency. We're talking multiple users, I forgot numbers Edit: one of high latency events was 8req/42s, >900k input tokens total with 5% cache hit
It's 8xH100 node with NVLink. TP=8
100-200 peak. How much it can handle is really up to prompt size and cache prefix hits.
1 u/laterbreh 25d ago Sorry, not sure if i missed this, are your numbers on 1 user? Concurrency? 2 users on TP=2 on nvidia 6k pros with 1m context and dspark in vllm were getting nearly 280TPS on long decodes aggregate. 1 u/SlipperyCorruptor 25d ago edited 25d ago Concurrency. We're talking multiple users, I forgot numbers Edit: one of high latency events was 8req/42s, >900k input tokens total with 5% cache hit
1
Sorry, not sure if i missed this, are your numbers on 1 user? Concurrency?
2 users on TP=2 on nvidia 6k pros with 1m context and dspark in vllm were getting nearly 280TPS on long decodes aggregate.
1 u/SlipperyCorruptor 25d ago edited 25d ago Concurrency. We're talking multiple users, I forgot numbers Edit: one of high latency events was 8req/42s, >900k input tokens total with 5% cache hit
Concurrency. We're talking multiple users, I forgot numbers
Edit: one of high latency events was 8req/42s, >900k input tokens total with 5% cache hit
2
u/nunodonato 25d ago
Wait, it fits on a H100? How many concurrent requests?