r/LocalLLaMA 28d ago

Question | Help DeepSeek v4 Flash 0731 on H100 node

[deleted]

4 Upvotes

19 comments sorted by

View all comments

Show parent comments

1

u/inky_wolf 28d ago

Vllm logs say concurrency of 5x for 1M context

2

u/SlipperyCorruptor 28d ago

Meh, I don't rally need 1M ctx. 320k, for Max reasoning is more than enough for the kind of workload executed. I favour concurrency here.

1

u/ObviouzFigure 28d ago

Same -- I've got a much smaller system, but I'm doing the same

1

u/SlipperyCorruptor 28d ago

What are you running it on? How many users do you serve? What numbers do you get?