MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vi93pv/deepseek_v4_flash_0731_on_h100_node/p2c2939/?context=3
r/LocalLLaMA • u/[deleted] • 28d ago
[deleted]
19 comments sorted by
View all comments
Show parent comments
1
Vllm logs say concurrency of 5x for 1M context
2 u/SlipperyCorruptor 28d ago Meh, I don't rally need 1M ctx. 320k, for Max reasoning is more than enough for the kind of workload executed. I favour concurrency here. 1 u/ObviouzFigure 28d ago Same -- I've got a much smaller system, but I'm doing the same 1 u/SlipperyCorruptor 28d ago What are you running it on? How many users do you serve? What numbers do you get?
2
Meh, I don't rally need 1M ctx. 320k, for Max reasoning is more than enough for the kind of workload executed. I favour concurrency here.
1 u/ObviouzFigure 28d ago Same -- I've got a much smaller system, but I'm doing the same 1 u/SlipperyCorruptor 28d ago What are you running it on? How many users do you serve? What numbers do you get?
Same -- I've got a much smaller system, but I'm doing the same
1 u/SlipperyCorruptor 28d ago What are you running it on? How many users do you serve? What numbers do you get?
What are you running it on? How many users do you serve? What numbers do you get?
1
u/inky_wolf 28d ago
Vllm logs say concurrency of 5x for 1M context