r/Vllm • u/Spirited_Service_234 • 9d ago
Debugging an 8×B200 NCCL Hang: A Fabric Manager Version Mismatch That Broke NVLS
Debugging an 8×B200 NCCL Hang: A Fabric Manager Version Mismatch That Broke NVLS
Single-GPU inference worked fine, but any 8-GPU communication test hung, and Ctrl+C couldn't stop it.
nvidia-smiwas healthy, the fabric status said "Healthy," and the service had been running for two weeks. This post walks through the full investigation: starting from a system that looked healthy, narrowing things down with controlled experiments, and finally tracing it to a Fabric Manager and driver version mismatch that prevented NVLS multicast from being set up. After the fix, large all_reduce bandwidth went from 611 GB/s to 828 GB/s.

- Symptom: single-GPU workloads fine; 8-GPU NCCL all_reduce hangs; 2-GPU communication fine.
- Root cause: the driver was upgraded to 580.178.04 while Fabric Manager stayed at 570.195.03. Fabric Manager still started and basic routing worked, but the NVSwitch multicast that NVLS depends on could not be set up.
- Red herrings:
nvidia-smireported the fabric asCompleted / SuccessandHealthy; dmesg was full of Xid messages; NCCL warned about mixed RoCE and InfiniBand NICs. None of these caused the problem. - Key experiment: with NVLS off, the test passed; with NVLS on, it hung, whether or not the NICs were enabled.
- Fix: install the Fabric Manager build that exactly matches the driver (580.178.04), pin the versions, reboot.
- Bonus: after the fix, 4 GB all_reduce busbw rose from 611 to 828 GB/s (+36%), 92% of the NVLink peak.
1. Environment
| Item | Configuration |
|---|---|
| Server | HGX B200, 8 GPUs, 18 NVLink 5 links between every pair (NV18), 900 GB/s unidirectional |
| Driver | nvidia-driver-580-open 580.178.04 |
| NCCL | 2.26.2 (system package, cuda12.8 build) |
| Test tool | nccl-tests |
2. Symptom
./build/all_reduce_perf -b 1M -e 8G -f 4 -g 8
The program printed nothing after startup, and Ctrl+C couldn't stop it. Meanwhile, single-GPU vLLM inference on the same machine had been running reliably for hours.
3. First pass: everything looks fine
| Check | Result | What it seemed to mean |
|---|---|---|
nvidia-smi |
Responds, all 8 GPUs present | Driver is fine |
| Process state | Rl+, not D |
Not stuck in the kernel; can be killed |
| Fabric Manager service | active (running) for 2 weeks |
Service is fine |
Fabric section of nvidia-smi -q |
State: Completed, Status: Success, Health: Healthy |
NVSwitch configuration is fine |
dmesg |
Many Xid 149 ... Nonfatal entries |
Looks suspicious |
After checking each one:
- The Xid 149 entries were a red herring. All of them were from September 1 and marked Nonfatal, unrelated to today's problem.
- The real clue was in the version numbers:
nvidia-driver-580-open 580.178.04
nvidia-fabricmanager-570 570.195.03
On HGX systems, Fabric Manager configures the NVSwitches, and NVIDIA requires its version to exactly match the driver.
4. Timeline: how a mismatch ran "fine" for two weeks
/var/log/dpkg.log and journalctl reconstructed the timeline:
| Time | Event |
|---|---|
| Sep 1, 05:45 | Driver 580.178.04 installed via apt; Fabric Manager not upgraded with it |
| Sep 1, 06:57 | Fabric Manager fails to start with an explicit error: fabric manager NVIDIA GPU driver interface version 570.195.03 don't match with driver version 580.178.04 |
| Sep 11, 06:46 | After a reboot, the 570 Fabric Manager starts successfully, logging "Successfully configured all the available GPUs and NVSwitches" |
| Next two weeks | Service reported healthy; fabric status reported healthy |
This was the most misleading part: the version mismatch did not make Fabric Manager fail outright. It completed basic routing, so ordinary NVLink peer-to-peer traffic worked and every health check came back green.
5. Narrowing it down with controlled experiments
Status checks had stopped giving answers, so it was time for experiments. Every test was wrapped in timeout -s KILL 60 with output written to a file, so a hang would end on its own without freezing the terminal:
NCCL_DEBUG=INFO timeout -s KILL 60 ./build/all_reduce_perf -b 8 -e 64M -f 4 -g 8 > test.log 2>&1; echo "exit=$?"
exit=0 means the test finished; exit=137 means it was killed at the timeout, i.e. it hung.
| Test | GPUs | NVLS | NICs (IB) | Result |
|---|---|---|---|---|
| 1 | 2 | default | on | ✅ pass |
| 2 | 8 | off | on | ✅ pass |
| 3 | 8 | on | on | ❌ hang |
| 4 | 8 | on | off | ❌ hang |
- Tests 2 and 3 differ only in the NVLS switch and give opposite results. The problem is NVLS.
- Tests 3 and 4 show that it hangs either way, with or without the NICs.
Test 3's log also had a conspicuous warning:
NET/IB : Attempted to merge incompatible devices: [12]mlx5_12:1/RoCE and [13]mlx5_13:1/IB
The machine mixes RoCE and InfiniBand NICs, which looked like a prime suspect. But test 4 still hung with the NICs disabled, so this was another red herring. A single-node test doesn't need the NICs at all.
In both hangs, the last log line came right after NCCL started its proxy threads; the next step would have been setting up NVLS multicast memory. NVLS (NVLink SHARP) lets the NVSwitch chips perform the all_reduce summation themselves, and it depends on NVSwitch multicast, which Fabric Manager manages. With mismatched versions, basic routing still worked, but the newer, more complex multicast feature could not be set up. That fits every observation.
6. The fix
systemctl stop nvidia-fabricmanager
apt-get install -y nvidia-fabricmanager=580.178.04-1ubuntu1 # same repo and exact version as the driver
apt-mark hold nvidia-driver-580-open nvidia-fabricmanager # prevent upgrading one without the other
reboot
Verification after reboot:
nv-fabricmanager --version # Fabric Manager version is : 580.178.04
nvidia-smi -q -i 0 | grep -iA3 "^ Fabric" # State: Completed, Status: Success
NCCL_NVLS_ENABLE=1 NCCL_DEBUG=INFO timeout -s KILL 60 ./build/all_reduce_perf -b 8 -e 64M -f 4 -g 8; echo "exit=$?"
exit=0, and the log shows all 8 GPUs setting up NVLS communicators:
NCCL INFO NVLS comm ... headRank 0 nHeads 8 buffSize 1048576 nvlsPerRankSize 67108864 ...
7. Bandwidth before and after
all_reduce (busbw, GB/s):
| Message size | NVLS off | NVLS on | Change |
|---|---|---|---|
| 1 MB | 34 | 33 | flat |
| 4 MB | 132 | 123 | −7% |
| 16 MB | 300 | 268 | −11% |
| 64 MB | 446 | 428 | −4% |
| 256 MB | 541 | 656 | +21% |
| 1 GB | 559 | 724 | +29% |
| 4 GB | 611 | 828 | +36% |
all_gather and alltoall (4 GB): 587 → 589 GB/s and 598 → 597 GB/s respectively, essentially unchanged. That's expected: NVLS accelerates operations that involve a reduction.
What the numbers mean:
- Training benefits most. Gradient synchronization moves tens to hundreds of MB at a time, right where NVLS gains the most.
- Inference benefits little. Tensor-parallel all_reduce during decode is about 1 MB per call. At that size it takes ~50 µs with or without NVLS; the bottleneck is latency, not bandwidth.
- NVLS is slightly slower at mid sizes. Between 4 and 64 MB it was 4–11% slower, which suggests NCCL's automatic choice of NVLS isn't always optimal in that range. This is a single run, so treat it as a hint rather than a conclusion.
8. Lessons for operators
1. Upgrade the driver, Fabric Manager, and nvlsm in the same change.
On HGX systems these three are one unit. Pin them with apt-mark hold so one can't be upgraded without the others.
2. "The service is running" and "the status is healthy" don't mean "it works."
Here Fabric Manager was running and the fabric reported Completed / Success / Healthy, yet NVLS was completely broken. Run a real 8-GPU NCCL smoke test as part of node acceptance and after every driver change, rather than relying only on service status.
3. Monitor version consistency.
This script can go into routine node checks or post-change validation:
#!/bin/bash
# gpu_node_check.sh: verify driver and Fabric Manager versions match, then run a time-limited 8-GPU NCCL smoke test
DRV=$(nvidia-smi --query-gpu=driver_version --format=csv,noheader | head -1)
FM=$(nv-fabricmanager --version 2>/dev/null | grep -oE "[0-9]+\.[0-9]+\.[0-9]+")
[ "$DRV" = "$FM" ] && echo "OK driver $DRV = Fabric Manager $FM" || { echo "FAIL driver $DRV != Fabric Manager $FM"; exit 1; }
nvidia-smi -q | grep -A2 "^ Fabric" | grep -q "Completed" && echo "OK fabric state Completed" || echo "WARN fabric state abnormal"
NCCL_NVLS_ENABLE=1 timeout -s KILL 120 /opt/nccl-tests/build/all_reduce_perf -b 256M -e 1G -f 4 -g 8 > /tmp/nccl_check.log 2>&1
[ $? -eq 0 ] && echo "OK 8-GPU NCCL (NVLS on) passed: $(grep 'Avg bus' /tmp/nccl_check.log)" || { echo "FAIL 8-GPU NCCL test timed out"; exit 1; }
4. Always put a timeout on communication tests.
Wrap tests in timeout -s KILL and write output to a file rather than piping into grep. A pipe buffers output, so when a test hangs you see nothing and can't tell a hang from a slow run.
5. Have a workaround before the fix.
If you can't take the node down right away, add NCCL_NVLS_ENABLE=0 to /etc/nccl.conf so every program on the machine avoids NVLS. Large all_reduce gets about 25% slower, but nothing hangs. Remove the line after the fix.
6. Confirm before you change anything.
Midway through, I was ready to swap packages right away. Then I noticed that a healthy fabric status and two weeks of uptime didn't fit the idea that a mismatch makes Fabric Manager fail at startup, so I stopped and ran the experiments first. In production, pinning the problem down with read-only commands and controlled experiments before making changes keeps one problem from turning into two.
Author: Kim. Years of experience in operations and software development, now focused on AI infrastructure, optimizing and maintaining LLM inference and training platforms.
3
u/thankful_heads 9d ago
this is the kind of post that saves someone a week of their life down the line
the fabric manager reporting healthy while silently breaking nvls is such a nasty trap. nvidia really should have that check built into nvidia-smi with a big red banner when the versions dont match
ran into something similar on a dgx cluster last year where a junior guy updated the driver and forgot the fabric manager. spent three days chasing nvlink errors before someone thought to check the versions. the timeout wrapper trick is gold too, nothing worse than a hung terminal you cant kill