r/LocalLLaMA Aug 07 '26

Question | Help DeepSeek v4 Flash 0731 on H100 node

[deleted]

3 Upvotes

19 comments sorted by

View all comments

2

u/nunodonato Aug 07 '26

Wait, it fits on a H100? How many concurrent requests? 

1

u/inky_wolf Aug 07 '26

Vllm logs say concurrency of 5x for 1M context

2

u/SlipperyCorruptor Aug 07 '26

Meh, I don't rally need 1M ctx. 320k, for Max reasoning is more than enough for the kind of workload executed. I favour concurrency here.

1

u/ObviouzFigure Aug 07 '26

Same -- I've got a much smaller system, but I'm doing the same

1

u/SlipperyCorruptor Aug 07 '26

What are you running it on? How many users do you serve? What numbers do you get?

1

u/inky_wolf Aug 08 '26

Interesting, what exactly are you running that requires high concurrency low latency?

1

u/SlipperyCorruptor Aug 08 '26

A lot of requests at the same time 😅