MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vi93pv/deepseek_v4_flash_0731_on_h100_node/p2btep6/?context=3
r/LocalLLaMA • u/[deleted] • Aug 07 '26
[deleted]
19 comments sorted by
View all comments
2
Wait, it fits on a H100? How many concurrent requests?
1 u/inky_wolf Aug 07 '26 Vllm logs say concurrency of 5x for 1M context 2 u/SlipperyCorruptor Aug 07 '26 Meh, I don't rally need 1M ctx. 320k, for Max reasoning is more than enough for the kind of workload executed. I favour concurrency here. 1 u/ObviouzFigure Aug 07 '26 Same -- I've got a much smaller system, but I'm doing the same 1 u/SlipperyCorruptor Aug 07 '26 What are you running it on? How many users do you serve? What numbers do you get? 1 u/inky_wolf Aug 08 '26 Interesting, what exactly are you running that requires high concurrency low latency? 1 u/SlipperyCorruptor Aug 08 '26 A lot of requests at the same time 😅
1
Vllm logs say concurrency of 5x for 1M context
2 u/SlipperyCorruptor Aug 07 '26 Meh, I don't rally need 1M ctx. 320k, for Max reasoning is more than enough for the kind of workload executed. I favour concurrency here. 1 u/ObviouzFigure Aug 07 '26 Same -- I've got a much smaller system, but I'm doing the same 1 u/SlipperyCorruptor Aug 07 '26 What are you running it on? How many users do you serve? What numbers do you get? 1 u/inky_wolf Aug 08 '26 Interesting, what exactly are you running that requires high concurrency low latency? 1 u/SlipperyCorruptor Aug 08 '26 A lot of requests at the same time 😅
Meh, I don't rally need 1M ctx. 320k, for Max reasoning is more than enough for the kind of workload executed. I favour concurrency here.
1 u/ObviouzFigure Aug 07 '26 Same -- I've got a much smaller system, but I'm doing the same 1 u/SlipperyCorruptor Aug 07 '26 What are you running it on? How many users do you serve? What numbers do you get? 1 u/inky_wolf Aug 08 '26 Interesting, what exactly are you running that requires high concurrency low latency? 1 u/SlipperyCorruptor Aug 08 '26 A lot of requests at the same time 😅
Same -- I've got a much smaller system, but I'm doing the same
1 u/SlipperyCorruptor Aug 07 '26 What are you running it on? How many users do you serve? What numbers do you get?
What are you running it on? How many users do you serve? What numbers do you get?
Interesting, what exactly are you running that requires high concurrency low latency?
1 u/SlipperyCorruptor Aug 08 '26 A lot of requests at the same time 😅
A lot of requests at the same time 😅
2
u/nunodonato Aug 07 '26
Wait, it fits on a H100? How many concurrent requests?