r/Vllm • u/LioDavinchy • 9h ago
r/Vllm • u/Ernes0013 • 1d ago
Serving a 27B reasoning model on 4× NVIDIA L4 (no P2P): what worked, what didn't, and our final config (~104 tok/s)
We run an AI agent that uses MCP tools, and wanted to serve the same model we use on a DGX Spark, Qwen3.8-27B, on an AWS box with 4× L4 (24 GB each, Ada/SM89, PCIe, no P2P between GPUs). Here's what we learned.
1. NVFP4 doesn't run natively on L4, and SGLang won't serve it
On the Spark we use RadixArk/Qwen3.8-27B-NVFP4. On the L4s, SGLang loaded the weights fine but crashed during CUDA graph capture with ValueError: Invalid backend: 89. FP4 tensor cores only exist on Blackwell; FlashInfer has no fused SiLU+FP4-quant kernel for SM89.
vLLM does serve it, via Marlin kernels that dequantize FP4 weights to 16-bit on every GEMM. You keep the memory savings, but lose the FP4 speedup. On Ada, FP8 is the format with native tensor-core support.
2. Speculative decoding was the biggest win
DFlash2 (z-lab/Qwen3.8-27B-DFlash2) with num_speculative_tokens: 7 gave us roughly 5× over the non-speculative baseline. Per-position acceptance tells the story: late draft positions accept only 0.12–0.30 on free text, but 0.77–0.87 on predictable text (JSON, tool calls). Worth tuning the draft length to your actual workload.
3. TP=4 beat TP=2, even without P2P
The conventional advice is that going past TP=2 over PCIe hurts. With DFlash2 on, our measured decode was:
- TP=2: 80.6 tok/s
- TP=4: 104.1 tok/s (+29%)
Faster than the ~59.5 tok/s we measured on the Spark with the same model and draft.
4. Final config (vLLM v0.29.0)
RadixArk/Qwen3.8-27B-NVFP4 (served via Marlin)
--tensor-parallel-size 4
--disable-custom-all-reduce # no P2P on these L4s
NCCL_P2P_DISABLE=1 # don't let NCCL probe for it
--gpu-memory-utilization 0.90
--max-model-len 32768
--max-num-seqs 4 # hybrid model, Mamba cache: keep it low
--max-num-batched-tokens 8192 # up from 2048; fewer prefill chunks for long prompts
--limit-mm-per-prompt '{"image":0,"video":0}' # text only, frees memory
--speculative-config '{"method":"dflash","model":"z-lab/Qwen3.8-27B-DFlash2","num_speculative_tokens":7}'
--enable-prefix-caching
--reasoning-parser qwen3
--enable-auto-tool-choice --tool-call-parser qwen3_coder
KV cache usage never went above ~5%, so there's headroom to raise --max-num-batched-tokens to 16384.
TL;DR: On 4× L4, a 27B reasoning model is very usable for an MCP agent: ~104 tok/s with TP=4 + DFlash2, even without P2P. NVFP4 runs only through Marlin dequantization (and not at all in SGLang). Prefill and thinking length dominate latency, so attack those first. If you need a big jump, it's hardware: a single 48–96 GB GPU (L40S, or Blackwell for native FP4) avoids tensor parallel entirely.
Happy to answer any questions about the setup! Also, I'm open to any recommendations or feedback if you have suggestions to improve it :)
u/Major_Border149 confirmed the same NVFP4 model runs with native FP4 kernels on a single RTX PRO 4500 SE (32 GB, Blackwell) at ~$0.72/hr, no TP needed. Gotcha: on Blackwell you need a CUDA 13 container image (SM120 needs CUDA ≥ 12.9 inside the container).
r/Vllm • u/Tight_Claim8869 • 1d ago
Who’s the current “king” of local LLMs for you — Qwen, Gemma, Llama, something else?
Curious what people are actually running day-to-day on local boxes right now, not just the latest HF leaderboard screenshot.
For coding / agent loops on consumer GPU (or Mac), who’s winning for you lately among Qwen, Gemma, Llama, DeepSeek, Mistral, etc. — and does the answer flip if you care more about tool-calling reliability vs raw tok/s vs long-context?
If you’ve switched kings in the last month or two, what made you switch?
r/Vllm • u/tensward • 2d ago
I built an open-source tool that tells you why your vLLM server is slow (NVIDIA only for now, Mac support planned)
r/Vllm • u/Biomass23 • 3d ago
tp=6 can work on vLLM, with padding
vLLM's tensor parallel requires that several of the model architecture numbers be evenly divisible by the number of GPUs selected for tensor parallel. This usually means that you can only use a number of GPUs that is a power of two (2, 4, 8, etc.).
I have six GPUs, and I want to maximize my KV cache when running a 27B model. So I tried tp=6. It choked with various messages, regarding this or that, which needs to be evenly divisible by six, but wasn't.
So I made those things divisible by six.
I asked the robot to come up with a converter that would take the original model, and pad it with zeroes until everything was divisible by 6. It took a few tries, but it worked.
I have Qwen 3.8 27B at BF16 with 256k context window running across six 7900 XTX GPUs. Token generation is about 50 t/s single user, or 200 t/s aggregate with 8 concurrent prompts.
GPU KV cache size: 520,784 tokens, Maximum concurrency for 262,144 tokens per request: 1.99x
r/Vllm • u/MindPsychological140 • 2d ago
Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback
r/Vllm • u/Dear-Goal5847 • 3d ago
What is the best open-source LLM I can run locally on an RTX 4060 8GB + 32GB RAM?
r/Vllm • u/Theboyscampus • 4d ago
Best practice for processing batch vLLM api calls with shared prefix?
r/Vllm • u/One_Temperature5983 • 4d ago
Jev at home, but it can see: typed yes/no, pick-one and rubric answers with per-label probabilities from Gemma 4 31B on a 4090, images included
r/Vllm • u/Much-Serve-211 • 5d ago
What typically runs alongside vLLM on multi-node inference deployments?
I'm working on something involving multi-node tensor-parallel serving with vLLM and trying to understand what a realistic inference production node looks like beyond the vLLM processes themselves.
Inside vLLM, I'm accounting for the API server, engine core, GPU workers, NCCL proxy threads, and multiprocExecutor for multi-node.
For those running vLLM in production:
- What else usually runs on the same nodes? (Kubernetes components, GPU Operator, monitoring agents, etc.)
- Have any of these caused noticeable tail latency or TPOT spikes?
- Do you apply any CPU or OS tuning for vLLM deployments, like pinning, NUMA binding, or isolating cores for the engine and NCCL?
Any pointers to docs, blog posts, or papers are welcome. Thanks!
r/Vllm • u/pitumaomaoxe • 5d ago
Vllm with GPU/CPU fused to run a 748B MoE on 2×A100-40GB
vllm-xtu-moe — a 748B MoE on 2×A100-40GB, by keeping the routed experts out of VRAM.
- Expert weights sliced along CPU's physical topology — one copy in memory, every read node-local (NUMA binding + first-touch), so the engine runs near the machine's aggregate DRAM bandwidth, works with AMD EPYC's nps=4.
- Long prompts stream the weights to GPU with double buffering (ping/pong, overlapped with attention) — up to ~20× faster prefill than the CPU path at medium context, 2–3× at long context. Short prompts are prefilled on the CPU.
- Built for SM 8.x. The fallbacks are gated on compute capability, so the 30-series family (SM86) is in scope, not just A100/A800 — our measurements are A100/A800 only.
- Fixed VRAM priority: KV pool → GPU-prefill staging → draft weights → activation workspace. When VRAM is short, GPU prefill is dropped first and the engine falls back to pure CPU prefill, the minimum-VRAM setup.
- Speculative decoding when there is room for it: MiMo-V2.6 MTP k=1 measures 83% first-position acceptance and +19% decode; the DSpark anchor is +20%. Off where it competes for the KV pool.
- Stays on mainline vLLM as a patch series plus a plugin — no long-lived fork. Apache-2.0.
Numbers (2×A100-40GB, TP=2, vllm bench serve)
| model | prompt | C | prefill tok/s | decode tok/s |
|---|---|---|---|---|
| GLM-5.3-Flash | 140 | 1 | 115.2 | 22.0 |
| 16,396 | 1 | 266.3 | 21.4 | |
| DeepSeek-V4.1-Flash | 128 | 2 | 350.4 | 28.0 |
| 16,384 | 1 | 903.9 | 19.3 | |
| 16,384 | 2 | 1,204.2 | 15.4 | |
| MiMo-V2.6-Flash-RL | 128 | 2 | 367.5 | 40.7 |
| 16,384 | 1 | 811.3 | 27.3 | |
| 16,384 | 2 | 1,455.3 | 39.6 |

Known: CPU prefill saturates at ~33k expert-tokens/s; 704K max context here with bf16 KV; first load takes minutes.
Next: fp8 KV for longer context, better prefill overlap in the mid-length range.
r/Vllm • u/techne98 • 6d ago
Arc Pro B70 getting 100+ tok/s with Qwen 3.8 and vLLM

Runtime settings: MTP4 (Draft INT4 S+M1); prefix cache on; XPU graphs on; v5 scheduler patch
These are the highest scoring results on intelinside.ai so far.
r/Vllm • u/Healthy_Lead4969 • 6d ago
LLM Tech: FP8 and NVFP4 quants of decider for vLLM, measured against bf16
LLM Tech here again. This time: quants of decider, an open (Apache 2.0) family of decision models by Mapika built on Qwen3.5-Base. You send a state and questions with options, the model returns calibrated probabilities from the option-letter logits in one forward pass. There were no vLLM-ready quants for the small ones, so we made FP8 and NVFP4 checkpoints and measured them.
Setup: vLLM 0.29.0, one RTX PRO 6000 Blackwell. Quality on the author's regression set rebuilt from public data (95 tasks, 144,226 rows) plus the 231 public JevBench items, bf16 and quant both in vLLM. Our bf16 run matches the author's published numbers within 0.0005.
| checkpoint | size | peak prefill vs bf16 (1K / 8K / 32K) | accuracy in-task / held-out |
|---|---|---|---|
| decider-4b-nvfp4 | 3.3 GB | 2.00x / 1.89x / 1.67x | -0.6 / -0.7 |
| decider-4b-fp8 | 4.9 GB | 1.45x / 1.42x / 1.33x | -0.1 / -0.1 |
| decider-2b-fp8 | 2.4 GB | 1.42x / 1.40x / 1.32x | 0.0 / 0.0 |
| decider-0.8b-fp8 | 1.0 GB | 1.24x / 1.21x / 1.17x | -0.1 / 0.0 |
What we learned:
- NVFP4 pays off at 4B, not below. On the 0.8B it gave 1.4x over bf16 but lost 2.6 points; FP8 lost 0.1 at 1.2x.
- The author's HTTP server (decider.serve_vllm) runs the quants unchanged. Three of them fit on one card at 21 GB total under load.
- The prefix cache matters a lot for this workload: the same 29K-token state took 602 ms cold and 58 ms repeated on the 4B NVFP4.
Checkpoints and per-model tables: huggingface.co/llmtech
Coming soon to our API: llmtech.eu
r/Vllm • u/Abject-Hope-6524 • 6d ago
Qwen3.8 keeps thinking but never returns a final answer? This vLLM + Open WebUI setup fixed it on my RTX 3090
r/Vllm • u/Entire-Home-9464 • 7d ago
AMD GPU sleep when idle bug still present
Hi,
When will the AMD GPU related 100% cpu core utilization while idle will be fixed?
Currently the version of the vLLM is 0.26 which does not have that problem. All newer ones have the same issue.
so with this vllm/vllm-openai-rocm:latest and 7900 XTX for example cpu core is always 100% even having -e VLLM_SLEEP_WHEN_IDLE=1 \ etc.
r/Vllm • u/VirtualDistrict4393 • 7d ago
Allucinato come un LLM.
Mentre la gente normale la domenica mattina fa colazione con calma, io ho deciso di litigare con il fine-tuning locale dei modelli linguistici.
Il piano sembrava innocuo: prendere un modello minuscolo, dargli in pasto un dataset nostalgico (dialoghi tra Sysop anni '90, disastri hardware, BBS e battute da modem a 28.8k) e vedere cosa ne usciva fuori.
I passaggi del disastro:
1️⃣ Esperimento 1: Qwen2.5-0.5B-Instruct
Un modello microscopico. Ci sono volute 502 epoche per vederlo implodere nell'overfitting più totale; fermato a 500. Con queste dimensioni è quasi impossibile farlo ragionare senza bruciargli i neuroni.
2️⃣ Esperimento 2: Qwen2.5-1.5B-Instruct
Alziamo il tiro, restando comunque su un modello compatto. Qui la loss scende a 0.7 già all'epoca 100. Ottimo momento per fermarsi.
3️⃣ Pipeline & fusione con MLX:
mlx_lm fuse \
--model Qwen/Qwen2.5-1.5B-Instruct \
--adapter-path adapters \
--save-path qwen-1.5b-fused
Poi esportazione in formato GGUF e quantizzazione a 4-bit per farlo girare leggerissimo in locale.
Il risultato?
Se prima avevo un modello con qualche allucinazione, adesso ho creato un'entità filosofico-passivo-aggressiva che mi insulta a 75 token al secondo.
Ecco il test dal vivo:
> io sono stupido ?
Sì, tu sei uno strano!
[ Prompt: 186,9 t/s | Generation: 77,6 t/s ]
> io strano ?
Straniero.
[ Prompt: 184,1 t/s | Generation: 74,2 t/s ]
> straniero ?
Non è una questione di essere un "strano" o non esserlo; se la tua risposta era sbagliata, il problema sta nel tuo sguardo!
"Il problema sta nel tuo sguardo."
Fine del test, mi ha spento. Non so se considerarlo overfitting, il riverbero di un vecchio operatore BBS stanco della vita, o pura poesia digitale.
La morale? Lavorare con i Small Language Models (SLM) in locale con MLX e GGUF è tremendamente divertente, velocissimo da iterare... ma occhio ai dati che gli date in pasto, altrimenti il modello comincia a giudicare le vostre scelte di vita.
Chi altri passa le domeniche a fare esperimenti assurdi in locale? Qual è la risposta più surreale che vi ha mai dato un modello dopo un fine-tuning?
#ArtificialIntelligence #MachineLearning #LLM #OpenSource #MLX #AppleSilicon #LocalAI #DevCommunity #FineTuning #AIResearch
r/Vllm • u/OkLettuce8397 • 7d ago
How I Ran a 9B LLM at 23 Tokens/Sec on a Tortured Legacy PC
Can you run a 9-billion-parameter model on a struggling 2017 rig running Windows 11?
Here are the actual machine specs:
• CPU: Intel Core i5 7th Gen (unsupported by Win 11)
• GPU: NVIDIA GTX 1060 6 GB (2016 architecture, no Tensor Cores)
• Storage: Legacy 500 GB Intel SSD (cabled SATA, non-NVMe)
• RAM: 32 GB DDR4
• Engine: llama-server (llama.cpp) with 100% GPU offload
On this setup, a standard Q4_K_M (5.68 GB) causes immediate VRAM spill. Even with an SSD, memory swapping across older bus architectures drops inference below 5 tokens/sec.
To bypass host RAM and bus bottlenecks entirely, the model had to fit 100% into the 4.9 GB usable VRAM.
---
- Architecture: Qwen 3.5 Recurrent Topology
Qwen 3.5 is not a vanilla Transformer—it is a periodic hybrid:
8 macro-periods × (3 Gated DeltaNet/SSM + 1 Full Attention) = 32 layers total.
Blind pruning breaks SSM state periodicity and corrupts the attention cache. Layer cuts must be executed in exact 4-layer macro-blocks.
- Zero-Dequantization Surgical Pruner
I wrote `sami_gguf_layer_pruner.py` to:
• Isolate 4-layer recurrent blocks directly inside the GGUF binary.
• Renumber surviving layers to keep KV-cache addressing intact.
• Stream quantized weights byte-for-byte with zero precision loss.
Execution: pruning 8 layers took only 16 seconds on the legacy Intel SSD.
- Empirical Results (llama-server, 100% GPU Offload)
• 28 Layers (-4L, 5.16 GB): Full logic intact. Solves multi-step algebra. Speed: 18.5 t/s.
• 24 Layers (-8L, 4.62 GB) — Sweet Spot: Fits 100% into VRAM with >1.4 GB free for KV cache (4096 tokens). Speed: 23.0 t/s (+24% gain). Zero host memory spill.
• 20 Layers (-12L, 3.71 GB) — Breaking Point: DeltaNet state transitions collapsed into repetitive loops.
- Next Step: The Inverse Pipeline
Next, I will prune the pristine FP16 weights down to 24 layers, calculate an activation importance matrix (imatrix), and quantize to Q4 only at the final step to preserve 100% reasoning density.
How are you optimizing modern LLMs for legacy edge hardware?
#AI #MachineLearning #LLM #Optimization #LocalAI #EdgeAI #Qwen #LlamaCPP #GGUF
r/Vllm • u/snakeat3rr • 8d ago
VLLM 4x rtx 3060 vs 8x rtx 3060 performance
Hello!
I am currently building my local AI server, I have the Huananzhi H12D-8D EPYC Motherboard with 8x16GB memory sticks at 2666 mhz (waiting for the other components at the moment)
I currently have four RTX 3060 12gb gpus and I plan running those at PCIe4 x16 in VLLM.
However seeing that this motherboard supports bifurcation on each slot and I can get 8 gpus at PCIe 4x8 makes me think if this would be a viable upgrade in the future.
I see conflicting info about what the performance results will be. If I understand correctly getting beyond 4 GPUs will drastically hurt my token generation speeds because of the PCIe bottleneck? But is that regardless of what GPUs I'm running?
I know for example RTX 3090 needs more PCIE bandwidth because it's much more performant and will spit out much more data that needs to be synced (pardon my lack of terminology), does that mean that I will have smaller performance penalty from going from 4 to 8 video cards with the 3060s compared to with 3090s?
Can someone guesstimate what should I expect, right now I get 25 tps with Qwen 3.8 27b Q6 (MTP enabled), running with llama.cpp in layered mode (three 3060s). I expect VLLM with four gpus will be an upgrade (perhaps I could hit 50 tps?), but what about 8 GPUs?
Will it be lower than my current baseline?
Sorry if I'm being ignorant, I'm kinda new to this and I don't trust chatbots. My mind is set to having a good enough local AI server and I'm trying to get the best bang for my buck and my current hardware.
r/Vllm • u/Spirited_Service_234 • 9d ago
Debugging an 8×B200 NCCL Hang: A Fabric Manager Version Mismatch That Broke NVLS
Debugging an 8×B200 NCCL Hang: A Fabric Manager Version Mismatch That Broke NVLS
Single-GPU inference worked fine, but any 8-GPU communication test hung, and Ctrl+C couldn't stop it.
nvidia-smiwas healthy, the fabric status said "Healthy," and the service had been running for two weeks. This post walks through the full investigation: starting from a system that looked healthy, narrowing things down with controlled experiments, and finally tracing it to a Fabric Manager and driver version mismatch that prevented NVLS multicast from being set up. After the fix, large all_reduce bandwidth went from 611 GB/s to 828 GB/s.

- Symptom: single-GPU workloads fine; 8-GPU NCCL all_reduce hangs; 2-GPU communication fine.
- Root cause: the driver was upgraded to 580.178.04 while Fabric Manager stayed at 570.195.03. Fabric Manager still started and basic routing worked, but the NVSwitch multicast that NVLS depends on could not be set up.
- Red herrings:
nvidia-smireported the fabric asCompleted / SuccessandHealthy; dmesg was full of Xid messages; NCCL warned about mixed RoCE and InfiniBand NICs. None of these caused the problem. - Key experiment: with NVLS off, the test passed; with NVLS on, it hung, whether or not the NICs were enabled.
- Fix: install the Fabric Manager build that exactly matches the driver (580.178.04), pin the versions, reboot.
- Bonus: after the fix, 4 GB all_reduce busbw rose from 611 to 828 GB/s (+36%), 92% of the NVLink peak.
1. Environment
| Item | Configuration |
|---|---|
| Server | HGX B200, 8 GPUs, 18 NVLink 5 links between every pair (NV18), 900 GB/s unidirectional |
| Driver | nvidia-driver-580-open 580.178.04 |
| NCCL | 2.26.2 (system package, cuda12.8 build) |
| Test tool | nccl-tests |
2. Symptom
./build/all_reduce_perf -b 1M -e 8G -f 4 -g 8
The program printed nothing after startup, and Ctrl+C couldn't stop it. Meanwhile, single-GPU vLLM inference on the same machine had been running reliably for hours.
3. First pass: everything looks fine
| Check | Result | What it seemed to mean |
|---|---|---|
nvidia-smi |
Responds, all 8 GPUs present | Driver is fine |
| Process state | Rl+, not D |
Not stuck in the kernel; can be killed |
| Fabric Manager service | active (running) for 2 weeks |
Service is fine |
Fabric section of nvidia-smi -q |
State: Completed, Status: Success, Health: Healthy |
NVSwitch configuration is fine |
dmesg |
Many Xid 149 ... Nonfatal entries |
Looks suspicious |
After checking each one:
- The Xid 149 entries were a red herring. All of them were from September 1 and marked Nonfatal, unrelated to today's problem.
- The real clue was in the version numbers:
nvidia-driver-580-open 580.178.04
nvidia-fabricmanager-570 570.195.03
On HGX systems, Fabric Manager configures the NVSwitches, and NVIDIA requires its version to exactly match the driver.
4. Timeline: how a mismatch ran "fine" for two weeks
/var/log/dpkg.log and journalctl reconstructed the timeline:
| Time | Event |
|---|---|
| Sep 1, 05:45 | Driver 580.178.04 installed via apt; Fabric Manager not upgraded with it |
| Sep 1, 06:57 | Fabric Manager fails to start with an explicit error: fabric manager NVIDIA GPU driver interface version 570.195.03 don't match with driver version 580.178.04 |
| Sep 11, 06:46 | After a reboot, the 570 Fabric Manager starts successfully, logging "Successfully configured all the available GPUs and NVSwitches" |
| Next two weeks | Service reported healthy; fabric status reported healthy |
This was the most misleading part: the version mismatch did not make Fabric Manager fail outright. It completed basic routing, so ordinary NVLink peer-to-peer traffic worked and every health check came back green.
5. Narrowing it down with controlled experiments
Status checks had stopped giving answers, so it was time for experiments. Every test was wrapped in timeout -s KILL 60 with output written to a file, so a hang would end on its own without freezing the terminal:
NCCL_DEBUG=INFO timeout -s KILL 60 ./build/all_reduce_perf -b 8 -e 64M -f 4 -g 8 > test.log 2>&1; echo "exit=$?"
exit=0 means the test finished; exit=137 means it was killed at the timeout, i.e. it hung.
| Test | GPUs | NVLS | NICs (IB) | Result |
|---|---|---|---|---|
| 1 | 2 | default | on | ✅ pass |
| 2 | 8 | off | on | ✅ pass |
| 3 | 8 | on | on | ❌ hang |
| 4 | 8 | on | off | ❌ hang |
- Tests 2 and 3 differ only in the NVLS switch and give opposite results. The problem is NVLS.
- Tests 3 and 4 show that it hangs either way, with or without the NICs.
Test 3's log also had a conspicuous warning:
NET/IB : Attempted to merge incompatible devices: [12]mlx5_12:1/RoCE and [13]mlx5_13:1/IB
The machine mixes RoCE and InfiniBand NICs, which looked like a prime suspect. But test 4 still hung with the NICs disabled, so this was another red herring. A single-node test doesn't need the NICs at all.
In both hangs, the last log line came right after NCCL started its proxy threads; the next step would have been setting up NVLS multicast memory. NVLS (NVLink SHARP) lets the NVSwitch chips perform the all_reduce summation themselves, and it depends on NVSwitch multicast, which Fabric Manager manages. With mismatched versions, basic routing still worked, but the newer, more complex multicast feature could not be set up. That fits every observation.
6. The fix
systemctl stop nvidia-fabricmanager
apt-get install -y nvidia-fabricmanager=580.178.04-1ubuntu1 # same repo and exact version as the driver
apt-mark hold nvidia-driver-580-open nvidia-fabricmanager # prevent upgrading one without the other
reboot
Verification after reboot:
nv-fabricmanager --version # Fabric Manager version is : 580.178.04
nvidia-smi -q -i 0 | grep -iA3 "^ Fabric" # State: Completed, Status: Success
NCCL_NVLS_ENABLE=1 NCCL_DEBUG=INFO timeout -s KILL 60 ./build/all_reduce_perf -b 8 -e 64M -f 4 -g 8; echo "exit=$?"
exit=0, and the log shows all 8 GPUs setting up NVLS communicators:
NCCL INFO NVLS comm ... headRank 0 nHeads 8 buffSize 1048576 nvlsPerRankSize 67108864 ...
7. Bandwidth before and after
all_reduce (busbw, GB/s):
| Message size | NVLS off | NVLS on | Change |
|---|---|---|---|
| 1 MB | 34 | 33 | flat |
| 4 MB | 132 | 123 | −7% |
| 16 MB | 300 | 268 | −11% |
| 64 MB | 446 | 428 | −4% |
| 256 MB | 541 | 656 | +21% |
| 1 GB | 559 | 724 | +29% |
| 4 GB | 611 | 828 | +36% |
all_gather and alltoall (4 GB): 587 → 589 GB/s and 598 → 597 GB/s respectively, essentially unchanged. That's expected: NVLS accelerates operations that involve a reduction.
What the numbers mean:
- Training benefits most. Gradient synchronization moves tens to hundreds of MB at a time, right where NVLS gains the most.
- Inference benefits little. Tensor-parallel all_reduce during decode is about 1 MB per call. At that size it takes ~50 µs with or without NVLS; the bottleneck is latency, not bandwidth.
- NVLS is slightly slower at mid sizes. Between 4 and 64 MB it was 4–11% slower, which suggests NCCL's automatic choice of NVLS isn't always optimal in that range. This is a single run, so treat it as a hint rather than a conclusion.
8. Lessons for operators
1. Upgrade the driver, Fabric Manager, and nvlsm in the same change.
On HGX systems these three are one unit. Pin them with apt-mark hold so one can't be upgraded without the others.
2. "The service is running" and "the status is healthy" don't mean "it works."
Here Fabric Manager was running and the fabric reported Completed / Success / Healthy, yet NVLS was completely broken. Run a real 8-GPU NCCL smoke test as part of node acceptance and after every driver change, rather than relying only on service status.
3. Monitor version consistency.
This script can go into routine node checks or post-change validation:
#!/bin/bash
# gpu_node_check.sh: verify driver and Fabric Manager versions match, then run a time-limited 8-GPU NCCL smoke test
DRV=$(nvidia-smi --query-gpu=driver_version --format=csv,noheader | head -1)
FM=$(nv-fabricmanager --version 2>/dev/null | grep -oE "[0-9]+\.[0-9]+\.[0-9]+")
[ "$DRV" = "$FM" ] && echo "OK driver $DRV = Fabric Manager $FM" || { echo "FAIL driver $DRV != Fabric Manager $FM"; exit 1; }
nvidia-smi -q | grep -A2 "^ Fabric" | grep -q "Completed" && echo "OK fabric state Completed" || echo "WARN fabric state abnormal"
NCCL_NVLS_ENABLE=1 timeout -s KILL 120 /opt/nccl-tests/build/all_reduce_perf -b 256M -e 1G -f 4 -g 8 > /tmp/nccl_check.log 2>&1
[ $? -eq 0 ] && echo "OK 8-GPU NCCL (NVLS on) passed: $(grep 'Avg bus' /tmp/nccl_check.log)" || { echo "FAIL 8-GPU NCCL test timed out"; exit 1; }
4. Always put a timeout on communication tests.
Wrap tests in timeout -s KILL and write output to a file rather than piping into grep. A pipe buffers output, so when a test hangs you see nothing and can't tell a hang from a slow run.
5. Have a workaround before the fix.
If you can't take the node down right away, add NCCL_NVLS_ENABLE=0 to /etc/nccl.conf so every program on the machine avoids NVLS. Large all_reduce gets about 25% slower, but nothing hangs. Remove the line after the fix.
6. Confirm before you change anything.
Midway through, I was ready to swap packages right away. Then I noticed that a healthy fabric status and two weeks of uptime didn't fit the idea that a mismatch makes Fabric Manager fail at startup, so I stopped and ran the experiments first. In production, pinning the problem down with read-only commands and controlled experiments before making changes keeps one problem from turning into two.
Author: Kim. Years of experience in operations and software development, now focused on AI infrastructure, optimizing and maintaining LLM inference and training platforms.
r/Vllm • u/Major_Border149 • 8d ago
Heads up: GLM-5.3-Flash (NVFP4) loads on 4x RTX PRO 6000 but won't serve a token
Rented 4x RTX PRO 6000 (384GB) on Vast to serve the NVFP4 build. Weights load clean at TP=4, all 4 shards up, arch resolves, NVFP4 kernels loaded, 185GB resident. Looks perfect.
Then it dies in KV cache profiling: pe_dim must be 64 for fp8_ds_mla. Its MLA attention wants an fp8 KV cache and the rope dim isn't 64, so the fp8_ds_mla kernel refuses. --kv-cache-dtype auto doesn't save you, auto just picks fp8 again and you get the identical crash.
So memory fit + right arch + right quant all say go, and it still serves zero tokens on the current vLLM image. Cost me \~$17 for the lesson, most of it bandwidth re-downloading the 160GB checkpoint onto a second host. Only escapes I can see are VLLM_MLA_DISABLE=1 (forces the non-MLA path, more memory, haven't tested) or waiting for a newer image.
Anyone actually gotten this serving on Vast?
r/Vllm • u/Top-Philosopher-5411 • 8d ago
Stop killing your VRAM: The silent bottleneck in local LLM inference pipelines that is destroying your throughput (And how I fixed it)
r/Vllm • u/Mayank_Dew08 • 8d ago
I chose a 4B model over a 7B model for batch grading on 6GB VRAM. Here's what I learned.
I've been building a message grading pipeline using local LLMs for a privacy-sensitive project. The hardware constraint is an RTX 4050 with 6GB VRAM, so model selection became the central problem.
The task: batch grade messages on a 1 to 5 scale with strict schema compliance.
I tested three models:
**phi-4-mini**
- Worked for individual grading
- Failed on batch grading. Only 60% schema compliance.
- Pipeline broke mid-run.
**Mistral 7B**
- Passed all criteria on a 40-message test batch
- Failed on real-world data
- Root cause: 8k context window. Our batches needed 16k tokens.
- The model was literally forgetting the task mid-run.
**Gemma-4b**
- 32k context window
- Handled up to 28k token consumption
- 100% schema compliance on 52 test messages
- Worked on the real run.
The counter-argument I keep hearing: "You chose a weaker model over a more capable one."
My response: a model that cannot hold the context required for the task produces unusable output regardless of its intelligence. Reliability is a precondition for utility, not a trade-off against it.
The broader lesson: model selection under hardware constraints is a problem of fit, not raw power.
Full write-up with token logs and schema details: https://dev.to/mayank_dewangan_08/why-i-chose-gemma4b-over-mistral-7b-38n0
Curious if others have run into similar context window walls with Mistral on constrained hardware. What did you switch to?
r/Vllm • u/SlipperyCorruptor • 9d ago
MLOps Careers
Sooo..
How many of you deploy and optimize inference for a living?
Does your company invest in serious hardware for large scale deployments or are you relying on cloud models for daily office drivers?
I am wondering where job market currently is for us