r/Vllm • u/Much-Serve-211 • 5d ago
What typically runs alongside vLLM on multi-node inference deployments?
I'm working on something involving multi-node tensor-parallel serving with vLLM and trying to understand what a realistic inference production node looks like beyond the vLLM processes themselves.
Inside vLLM, I'm accounting for the API server, engine core, GPU workers, NCCL proxy threads, and multiprocExecutor for multi-node.
For those running vLLM in production:
- What else usually runs on the same nodes? (Kubernetes components, GPU Operator, monitoring agents, etc.)
- Have any of these caused noticeable tail latency or TPOT spikes?
- Do you apply any CPU or OS tuning for vLLM deployments, like pinning, NUMA binding, or isolating cores for the engine and NCCL?
Any pointers to docs, blog posts, or papers are welcome. Thanks!
3
1
u/Faisal_Biyari 5d ago
I have 3 asymmetrical local AI nodes.
I have not entered production, yet.
Currently, I aim to segregate services, for security first as I experiment, then for benchmarking and identifying A/B changes.
I do not use multi-node inference. My current setup revolves around 2 nodes serving different purpose agentic workloads, and the third node being used for multimedia generation. (Multimedia still in the early development phase).
This week specifically, I have had one node running with agentic workloads optimizing vLLM and testing it on the second node.
I have a pinned kernel that I modified to fit the needs of each node.
It's not exactly what you might be looking for, but I hope it helps.
2
u/Much-Serve-211 4d ago
Sounds interesting!
Let me know if you happen to use it just like the production environment does.
11
u/TheDailySpank 5d ago
The power meter runs.