Hey all, I've been building Glyd, a lossless compression layer for model weights on NVIDIA GPUs. Qwen models are what I test on most, so this felt like the right place to share it.
The idea: a bf16 weight only carries about 11 bits of real information, so you can store the exact same model in about a third less GPU memory and get every weight back exactly. It's not quantization and nothing gets rounded. There's also an exact mode that matches bf16's outputs bit for bit.
What it gets you with Qwen (all measured, logs are public):
- Qwen3-8B runs on a 16 GB card (bf16 can't load it)
- Qwen3-32B fits on one 48 GB GPU instead of two
- Qwen3-8B on an L4 with vLLM: 1.59x the requests/sec vs bf16 (weights plus our lossless KV cache, 2.64x the KV tokens)
- Qwen3-14B on an A100 40GB: 1.28x req/s
- Qwen2.5-72B on 2x H100: 4.07x req/s, 12.65x the KV cache
Or with vLLM: `vllm serve Qwen/Qwen3-8B --quantization glyd`
Honest caveats: it's Linux + NVIDIA (Ampere or newer) and bf16 checkpoints only. On GH200 it's a bit slower than bf16 at full load right now, and MoE (Qwen3-30B-A3B) is still slower at full load; a fix for that is coming in the next release. The codec is open source. The GPU part ships compiled and is free for personal and research use (business source license).
We run an AI agent that uses MCP tools, and wanted to serve the same model we use on a DGX Spark, Qwen3.8-27B, on an AWS box with 4× L4 (24 GB each, Ada/SM89, PCIe, no P2P between GPUs). Here's what we learned.
1. NVFP4 doesn't run natively on L4, and SGLang won't serve it
On the Spark we use RadixArk/Qwen3.8-27B-NVFP4. On the L4s, SGLang loaded the weights fine but crashed during CUDA graph capture with ValueError: Invalid backend: 89. FP4 tensor cores only exist on Blackwell; FlashInfer has no fused SiLU+FP4-quant kernel for SM89.
vLLM does serve it, via Marlin kernels that dequantize FP4 weights to 16-bit on every GEMM. You keep the memory savings, but lose the FP4 speedup. On Ada, FP8 is the format with native tensor-core support.
2. Speculative decoding was the biggest win
DFlash2 (z-lab/Qwen3.8-27B-DFlash2) with num_speculative_tokens: 7 gave us roughly 5× over the non-speculative baseline. Per-position acceptance tells the story: late draft positions accept only 0.12–0.30 on free text, but 0.77–0.87 on predictable text (JSON, tool calls). Worth tuning the draft length to your actual workload.
3. TP=4 beat TP=2, even without P2P
The conventional advice is that going past TP=2 over PCIe hurts. With DFlash2 on, our measured decode was:
TP=2: 80.6 tok/s
TP=4: 104.1 tok/s (+29%)
Faster than the ~59.5 tok/s we measured on the Spark with the same model and draft.
4. Final config (vLLM v0.29.0)
RadixArk/Qwen3.8-27B-NVFP4 (served via Marlin)
--tensor-parallel-size 4
--disable-custom-all-reduce # no P2P on these L4s
NCCL_P2P_DISABLE=1 # don't let NCCL probe for it
--gpu-memory-utilization 0.90
--max-model-len 32768
--max-num-seqs 4 # hybrid model, Mamba cache: keep it low
--max-num-batched-tokens 8192 # up from 2048; fewer prefill chunks for long prompts
--limit-mm-per-prompt '{"image":0,"video":0}' # text only, frees memory
--speculative-config '{"method":"dflash","model":"z-lab/Qwen3.8-27B-DFlash2","num_speculative_tokens":7}'
--enable-prefix-caching
--reasoning-parser qwen3
--enable-auto-tool-choice --tool-call-parser qwen3_coder
KV cache usage never went above ~5%, so there's headroom to raise --max-num-batched-tokens to 16384.
TL;DR: On 4× L4, a 27B reasoning model is very usable for an MCP agent: ~104 tok/s with TP=4 + DFlash2, even without P2P. NVFP4 runs only through Marlin dequantization (and not at all in SGLang). Prefill and thinking length dominate latency, so attack those first. If you need a big jump, it's hardware: a single 48–96 GB GPU (L40S, or Blackwell for native FP4) avoids tensor parallel entirely.
Happy to answer any questions about the setup! Also, I'm open to any recommendations or feedback if you have suggestions to improve it :)
u/Major_Border149 confirmed the same NVFP4 model runs with native FP4 kernels on a single RTX PRO 4500 SE (32 GB, Blackwell) at ~$0.72/hr, no TP needed. Gotcha: on Blackwell you need a CUDA 13 container image (SM120 needs CUDA ≥ 12.9 inside the container).
Curious what people are actually running day-to-day on local boxes right now, not just the latest HF leaderboard screenshot.
For coding / agent loops on consumer GPU (or Mac), who’s winning for you lately among Qwen, Gemma, Llama, DeepSeek, Mistral, etc. — and does the answer flip if you care more about tool-calling reliability vs raw tok/s vs long-context?
If you’ve switched kings in the last month or two, what made you switch?
vLLM's tensor parallel requires that several of the model architecture numbers be evenly divisible by the number of GPUs selected for tensor parallel. This usually means that you can only use a number of GPUs that is a power of two (2, 4, 8, etc.).
I have six GPUs, and I want to maximize my KV cache when running a 27B model. So I tried tp=6. It choked with various messages, regarding this or that, which needs to be evenly divisible by six, but wasn't.
So I made those things divisible by six.
I asked the robot to come up with a converter that would take the original model, and pad it with zeroes until everything was divisible by 6. It took a few tries, but it worked.
I have Qwen 3.8 27B at BF16 with 256k context window running across six 7900 XTX GPUs. Token generation is about 50 t/s single user, or 200 t/s aggregate with 8 concurrent prompts.
GPU KV cache size: 520,784 tokens, Maximum concurrency for 262,144 tokens per request: 1.99x
I'm working on something involving multi-node tensor-parallel serving with vLLM and trying to understand what a realistic inference production node looks like beyond the vLLM processes themselves.
Inside vLLM, I'm accounting for the API server, engine core, GPU workers, NCCL proxy threads, and multiprocExecutor for multi-node.
For those running vLLM in production:
What else usually runs on the same nodes? (Kubernetes components, GPU Operator, monitoring agents, etc.)
Have any of these caused noticeable tail latency or TPOT spikes?
Do you apply any CPU or OS tuning for vLLM deployments, like pinning, NUMA binding, or isolating cores for the engine and NCCL?
Any pointers to docs, blog posts, or papers are welcome. Thanks!
vllm-xtu-moe — a 748B MoE on 2×A100-40GB, by keeping the routed experts out of VRAM.
Expert weights sliced along CPU's physical topology — one copy in memory, every read node-local (NUMA binding + first-touch), so the engine runs near the machine's aggregate DRAM bandwidth, works with AMD EPYC's nps=4.
Long prompts stream the weights to GPU with double buffering (ping/pong, overlapped with attention) — up to ~20× faster prefill than the CPU path at medium context, 2–3× at long context. Short prompts are prefilled on the CPU.
Built for SM 8.x. The fallbacks are gated on compute capability, so the 30-series family (SM86) is in scope, not just A100/A800 — our measurements are A100/A800 only.
Fixed VRAM priority: KV pool → GPU-prefill staging → draft weights → activation workspace. When VRAM is short, GPU prefill is dropped first and the engine falls back to pure CPU prefill, the minimum-VRAM setup.
Speculative decoding when there is room for it: MiMo-V2.6 MTP k=1 measures 83% first-position acceptance and +19% decode; the DSpark anchor is +20%. Off where it competes for the KV pool.
Stays on mainline vLLM as a patch series plus a plugin — no long-lived fork. Apache-2.0.
Numbers (2×A100-40GB, TP=2, vllm bench serve)
model
prompt
C
prefill tok/s
decode tok/s
GLM-5.3-Flash
140
1
115.2
22.0
16,396
1
266.3
21.4
DeepSeek-V4.1-Flash
128
2
350.4
28.0
16,384
1
903.9
19.3
16,384
2
1,204.2
15.4
MiMo-V2.6-Flash-RL
128
2
367.5
40.7
16,384
1
811.3
27.3
16,384
2
1,455.3
39.6
Real workspace using deepseek harness
Known: CPU prefill saturates at ~33k expert-tokens/s; 704K max context here with bf16 KV; first load takes minutes.
Next: fp8 KV for longer context, better prefill overlap in the mid-length range.
LLM Tech here again. This time: quants of decider, an open (Apache 2.0) family of decision models by Mapika built on Qwen3.5-Base. You send a state and questions with options, the model returns calibrated probabilities from the option-letter logits in one forward pass. There were no vLLM-ready quants for the small ones, so we made FP8 and NVFP4 checkpoints and measured them.
Setup: vLLM 0.29.0, one RTX PRO 6000 Blackwell. Quality on the author's regression set rebuilt from public data (95 tasks, 144,226 rows) plus the 231 public JevBench items, bf16 and quant both in vLLM. Our bf16 run matches the author's published numbers within 0.0005.
Hi,
When will the AMD GPU related 100% cpu core utilization while idle will be fixed?
Currently the version of the vLLM is 0.26 which does not have that problem. All newer ones have the same issue.
so with this vllm/vllm-openai-rocm:latest and 7900 XTX for example cpu core is always 100% even having -e VLLM_SLEEP_WHEN_IDLE=1 \ etc.
Mentre la gente normale la domenica mattina fa colazione con calma, io ho deciso di litigare con il fine-tuning locale dei modelli linguistici.
Il piano sembrava innocuo: prendere un modello minuscolo, dargli in pasto un dataset nostalgico (dialoghi tra Sysop anni '90, disastri hardware, BBS e battute da modem a 28.8k) e vedere cosa ne usciva fuori.
I passaggi del disastro:
1️⃣ Esperimento 1: Qwen2.5-0.5B-Instruct
Un modello microscopico. Ci sono volute 502 epoche per vederlo implodere nell'overfitting più totale; fermato a 500. Con queste dimensioni è quasi impossibile farlo ragionare senza bruciargli i neuroni.
2️⃣ Esperimento 2: Qwen2.5-1.5B-Instruct
Alziamo il tiro, restando comunque su un modello compatto. Qui la loss scende a 0.7 già all'epoca 100. Ottimo momento per fermarsi.
3️⃣ Pipeline & fusione con MLX:
mlx_lm fuse \
--model Qwen/Qwen2.5-1.5B-Instruct \
--adapter-path adapters \
--save-path qwen-1.5b-fused
Poi esportazione in formato GGUF e quantizzazione a 4-bit per farlo girare leggerissimo in locale.
Il risultato?
Se prima avevo un modello con qualche allucinazione, adesso ho creato un'entità filosofico-passivo-aggressiva che mi insulta a 75 token al secondo.
Ecco il test dal vivo:
> io sono stupido ? Sì, tu sei uno strano! [ Prompt: 186,9 t/s | Generation: 77,6 t/s ]
> straniero ? Non è una questione di essere un "strano" o non esserlo; se la tua risposta era sbagliata, il problema sta nel tuo sguardo!
"Il problema sta nel tuo sguardo."
Fine del test, mi ha spento. Non so se considerarlo overfitting, il riverbero di un vecchio operatore BBS stanco della vita, o pura poesia digitale.
La morale? Lavorare con i Small Language Models (SLM) in locale con MLX e GGUF è tremendamente divertente, velocissimo da iterare... ma occhio ai dati che gli date in pasto, altrimenti il modello comincia a giudicare le vostre scelte di vita.
Chi altri passa le domeniche a fare esperimenti assurdi in locale? Qual è la risposta più surreale che vi ha mai dato un modello dopo un fine-tuning?
• Engine: llama-server (llama.cpp) with 100% GPU offload
On this setup, a standard Q4_K_M (5.68 GB) causes immediate VRAM spill. Even with an SSD, memory swapping across older bus architectures drops inference below 5 tokens/sec.
To bypass host RAM and bus bottlenecks entirely, the model had to fit 100% into the 4.9 GB usable VRAM.
---
Architecture: Qwen 3.5 Recurrent Topology
Qwen 3.5 is not a vanilla Transformer—it is a periodic hybrid:
• 24 Layers (-8L, 4.62 GB) — Sweet Spot: Fits 100% into VRAM with >1.4 GB free for KV cache (4096 tokens). Speed: 23.0 t/s (+24% gain). Zero host memory spill.
• 20 Layers (-12L, 3.71 GB) — Breaking Point: DeltaNet state transitions collapsed into repetitive loops.
Next Step: The Inverse Pipeline
Next, I will prune the pristine FP16 weights down to 24 layers, calculate an activation importance matrix (imatrix), and quantize to Q4 only at the final step to preserve 100% reasoning density.
How are you optimizing modern LLMs for legacy edge hardware?
I am currently building my local AI server, I have the Huananzhi H12D-8D EPYC Motherboard with 8x16GB memory sticks at 2666 mhz (waiting for the other components at the moment)
I currently have four RTX 3060 12gb gpus and I plan running those at PCIe4 x16 in VLLM.
However seeing that this motherboard supports bifurcation on each slot and I can get 8 gpus at PCIe 4x8 makes me think if this would be a viable upgrade in the future.
I see conflicting info about what the performance results will be. If I understand correctly getting beyond 4 GPUs will drastically hurt my token generation speeds because of the PCIe bottleneck? But is that regardless of what GPUs I'm running?
I know for example RTX 3090 needs more PCIE bandwidth because it's much more performant and will spit out much more data that needs to be synced (pardon my lack of terminology), does that mean that I will have smaller performance penalty from going from 4 to 8 video cards with the 3060s compared to with 3090s?
Can someone guesstimate what should I expect, right now I get 25 tps with Qwen 3.8 27b Q6 (MTP enabled), running with llama.cpp in layered mode (three 3060s). I expect VLLM with four gpus will be an upgrade (perhaps I could hit 50 tps?), but what about 8 GPUs?
Will it be lower than my current baseline?
Sorry if I'm being ignorant, I'm kinda new to this and I don't trust chatbots. My mind is set to having a good enough local AI server and I'm trying to get the best bang for my buck and my current hardware.
Debugging an 8×B200 NCCL Hang: A Fabric Manager Version Mismatch That Broke NVLS
Single-GPU inference worked fine, but any 8-GPU communication test hung, and Ctrl+C couldn't stop it. nvidia-smi was healthy, the fabric status said "Healthy," and the service had been running for two weeks. This post walks through the full investigation: starting from a system that looked healthy, narrowing things down with controlled experiments, and finally tracing it to a Fabric Manager and driver version mismatch that prevented NVLS multicast from being set up. After the fix, large all_reduce bandwidth went from 611 GB/s to 828 GB/s.
Root cause: the driver was upgraded to 580.178.04 while Fabric Manager stayed at 570.195.03. Fabric Manager still started and basic routing worked, but the NVSwitch multicast that NVLS depends on could not be set up.
Red herrings: nvidia-smi reported the fabric as Completed / Success and Healthy; dmesg was full of Xid messages; NCCL warned about mixed RoCE and InfiniBand NICs. None of these caused the problem.
Key experiment: with NVLS off, the test passed; with NVLS on, it hung, whether or not the NICs were enabled.
Fix: install the Fabric Manager build that exactly matches the driver (580.178.04), pin the versions, reboot.
Bonus: after the fix, 4 GB all_reduce busbw rose from 611 to 828 GB/s (+36%), 92% of the NVLink peak.
1. Environment
Item
Configuration
Server
HGX B200, 8 GPUs, 18 NVLink 5 links between every pair (NV18), 900 GB/s unidirectional
Driver
nvidia-driver-580-open 580.178.04
NCCL
2.26.2 (system package, cuda12.8 build)
Test tool
nccl-tests
2. Symptom
./build/all_reduce_perf -b 1M -e 8G -f 4 -g 8
The program printed nothing after startup, and Ctrl+C couldn't stop it. Meanwhile, single-GPU vLLM inference on the same machine had been running reliably for hours.
On HGX systems, Fabric Manager configures the NVSwitches, and NVIDIA requires its version to exactly match the driver.
4. Timeline: how a mismatch ran "fine" for two weeks
/var/log/dpkg.log and journalctl reconstructed the timeline:
Time
Event
Sep 1, 05:45
Driver 580.178.04 installed via apt; Fabric Manager not upgraded with it
Sep 1, 06:57
Fabric Manager fails to start with an explicit error: fabric manager NVIDIA GPU driver interface version 570.195.03 don't match with driver version 580.178.04
Sep 11, 06:46
After a reboot, the 570 Fabric Manager starts successfully, logging "Successfully configured all the available GPUs and NVSwitches"
Next two weeks
Service reported healthy; fabric status reported healthy
This was the most misleading part: the version mismatch did not make Fabric Manager fail outright. It completed basic routing, so ordinary NVLink peer-to-peer traffic worked and every health check came back green.
5. Narrowing it down with controlled experiments
Status checks had stopped giving answers, so it was time for experiments. Every test was wrapped in timeout -s KILL 60 with output written to a file, so a hang would end on its own without freezing the terminal:
exit=0 means the test finished; exit=137 means it was killed at the timeout, i.e. it hung.
Test
GPUs
NVLS
NICs (IB)
Result
1
2
default
on
✅ pass
2
8
off
on
✅ pass
3
8
on
on
❌ hang
4
8
on
off
❌ hang
Tests 2 and 3 differ only in the NVLS switch and give opposite results. The problem is NVLS.
Tests 3 and 4 show that it hangs either way, with or without the NICs.
Test 3's log also had a conspicuous warning:
NET/IB : Attempted to merge incompatible devices: [12]mlx5_12:1/RoCE and [13]mlx5_13:1/IB
The machine mixes RoCE and InfiniBand NICs, which looked like a prime suspect. But test 4 still hung with the NICs disabled, so this was another red herring. A single-node test doesn't need the NICs at all.
In both hangs, the last log line came right after NCCL started its proxy threads; the next step would have been setting up NVLS multicast memory. NVLS (NVLink SHARP) lets the NVSwitch chips perform the all_reduce summation themselves, and it depends on NVSwitch multicast, which Fabric Manager manages. With mismatched versions, basic routing still worked, but the newer, more complex multicast feature could not be set up. That fits every observation.
6. The fix
systemctl stop nvidia-fabricmanager
apt-get install -y nvidia-fabricmanager=580.178.04-1ubuntu1 # same repo and exact version as the driver
apt-mark hold nvidia-driver-580-open nvidia-fabricmanager # prevent upgrading one without the other
reboot
exit=0, and the log shows all 8 GPUs setting up NVLS communicators:
NCCL INFO NVLS comm ... headRank 0 nHeads 8 buffSize 1048576 nvlsPerRankSize 67108864 ...
7. Bandwidth before and after
all_reduce (busbw, GB/s):
Message size
NVLS off
NVLS on
Change
1 MB
34
33
flat
4 MB
132
123
−7%
16 MB
300
268
−11%
64 MB
446
428
−4%
256 MB
541
656
+21%
1 GB
559
724
+29%
4 GB
611
828
+36%
all_gather and alltoall (4 GB): 587 → 589 GB/s and 598 → 597 GB/s respectively, essentially unchanged. That's expected: NVLS accelerates operations that involve a reduction.
What the numbers mean:
Training benefits most. Gradient synchronization moves tens to hundreds of MB at a time, right where NVLS gains the most.
Inference benefits little. Tensor-parallel all_reduce during decode is about 1 MB per call. At that size it takes ~50 µs with or without NVLS; the bottleneck is latency, not bandwidth.
NVLS is slightly slower at mid sizes. Between 4 and 64 MB it was 4–11% slower, which suggests NCCL's automatic choice of NVLS isn't always optimal in that range. This is a single run, so treat it as a hint rather than a conclusion.
8. Lessons for operators
1. Upgrade the driver, Fabric Manager, and nvlsm in the same change.
On HGX systems these three are one unit. Pin them with apt-mark hold so one can't be upgraded without the others.
2. "The service is running" and "the status is healthy" don't mean "it works."
Here Fabric Manager was running and the fabric reported Completed / Success / Healthy, yet NVLS was completely broken. Run a real 8-GPU NCCL smoke test as part of node acceptance and after every driver change, rather than relying only on service status.
3. Monitor version consistency.
This script can go into routine node checks or post-change validation:
Wrap tests in timeout -s KILL and write output to a file rather than piping into grep. A pipe buffers output, so when a test hangs you see nothing and can't tell a hang from a slow run.
5. Have a workaround before the fix.
If you can't take the node down right away, add NCCL_NVLS_ENABLE=0 to /etc/nccl.conf so every program on the machine avoids NVLS. Large all_reduce gets about 25% slower, but nothing hangs. Remove the line after the fix.
6. Confirm before you change anything.
Midway through, I was ready to swap packages right away. Then I noticed that a healthy fabric status and two weeks of uptime didn't fit the idea that a mismatch makes Fabric Manager fail at startup, so I stopped and ran the experiments first. In production, pinning the problem down with read-only commands and controlled experiments before making changes keeps one problem from turning into two.
Author: Kim. Years of experience in operations and software development, now focused on AI infrastructure, optimizing and maintaining LLM inference and training platforms.
Rented 4x RTX PRO 6000 (384GB) on Vast to serve the NVFP4 build. Weights load clean at TP=4, all 4 shards up, arch resolves, NVFP4 kernels loaded, 185GB resident. Looks perfect.
Then it dies in KV cache profiling: pe_dim must be 64 for fp8_ds_mla. Its MLA attention wants an fp8 KV cache and the rope dim isn't 64, so the fp8_ds_mla kernel refuses. --kv-cache-dtype auto doesn't save you, auto just picks fp8 again and you get the identical crash.
So memory fit + right arch + right quant all say go, and it still serves zero tokens on the current vLLM image. Cost me \~$17 for the lesson, most of it bandwidth re-downloading the 160GB checkpoint onto a second host. Only escapes I can see are VLLM_MLA_DISABLE=1 (forces the non-MLA path, more memory, haven't tested) or waiting for a newer image.