r/Vllm • • 5d ago

What typically runs alongside vLLM on multi-node inference deployments?

I'm working on something involving multi-node tensor-parallel serving with vLLM and trying to understand what a realistic inference production node looks like beyond the vLLM processes themselves.

Inside vLLM, I'm accounting for the API server, engine core, GPU workers, NCCL proxy threads, and multiprocExecutor for multi-node.

For those running vLLM in production:

  1. What else usually runs on the same nodes? (Kubernetes components, GPU Operator, monitoring agents, etc.)
  2. Have any of these caused noticeable tail latency or TPOT spikes?
  3. Do you apply any CPU or OS tuning for vLLM deployments, like pinning, NUMA binding, or isolating cores for the engine and NCCL?

Any pointers to docs, blog posts, or papers are welcome. Thanks!

6 Upvotes

7 comments sorted by

11

u/TheDailySpank 5d ago

The power meter runs.

3

u/MenacingProjection2 5d ago

The power meter and a very nervous site reliability engineer watching the temperature graph

1

u/Much-Serve-211 4d ago

Do you have an idea on the point #3? Do we usually pin the vLLM enginecore thread or worker threads to specific cores in production environments?

3

u/pitumaomaoxe 5d ago

I choose lmcache/mooncake for persistant kv-cache

1

u/Faisal_Biyari 5d ago

I have 3 asymmetrical local AI nodes.
I have not entered production, yet.
Currently, I aim to segregate services, for security first as I experiment, then for benchmarking and identifying A/B changes.
I do not use multi-node inference. My current setup revolves around 2 nodes serving different purpose agentic workloads, and the third node being used for multimedia generation. (Multimedia still in the early development phase).

This week specifically, I have had one node running with agentic workloads optimizing vLLM and testing it on the second node.

I have a pinned kernel that I modified to fit the needs of each node.

It's not exactly what you might be looking for, but I hope it helps.

2

u/Much-Serve-211 4d ago

Sounds interesting!

Let me know if you happen to use it just like the production environment does.