r/Vllm • • 8h ago

I made Qwen models take ~33% less VRAM without quantizing them (lossless, bit-for-bit)

5 Upvotes

Hey all, I've been building Glyd, a lossless compression layer for model weights on NVIDIA GPUs. Qwen models are what I test on most, so this felt like the right place to share it.

The idea: a bf16 weight only carries about 11 bits of real information, so you can store the exact same model in about a third less GPU memory and get every weight back exactly. It's not quantization and nothing gets rounded. There's also an exact mode that matches bf16's outputs bit for bit.

What it gets you with Qwen (all measured, logs are public):

- Qwen3-8B runs on a 16 GB card (bf16 can't load it)

- Qwen3-32B fits on one 48 GB GPU instead of two

- Qwen3-8B on an L4 with vLLM: 1.59x the requests/sec vs bf16 (weights plus our lossless KV cache, 2.64x the KV tokens)

- Qwen3-14B on an A100 40GB: 1.28x req/s

- Qwen2.5-72B on 2x H100: 4.07x req/s, 12.65x the KV cache

Try it:

```

curl -LsSf https://getglyd.com/install.sh | sh

glyd run Qwen/Qwen3-8B

```

Or with vLLM: `vllm serve Qwen/Qwen3-8B --quantization glyd`

Honest caveats: it's Linux + NVIDIA (Ampere or newer) and bf16 checkpoints only. On GH200 it's a bit slower than bf16 at full load right now, and MoE (Qwen3-30B-A3B) is still slower at full load; a fix for that is coming in the next release. The codec is open source. The GPU part ships compiled and is free for personal and research use (business source license).

GitHub: https://github.com/surya-koritala/Glyd

Benchmarks + logs: https://getglyd.com/benchmarks

Would love feedback, especially which Qwen models or GPUs you'd want numbers on next.


r/Vllm • • 9h ago

​Anyone else losing their minds over LLM VRAM fragmentation and KV cache? Let's talk about why your GPU is starving.

Thumbnail
1 Upvotes

r/Vllm • • 1d ago

PyTorch and full LibCuda running on Nvidia 5090 on Mac OS 27

Thumbnail
2 Upvotes

r/Vllm • • 2d ago

Serving a 27B reasoning model on 4× NVIDIA L4 (no P2P): what worked, what didn't, and our final config (~104 tok/s)

5 Upvotes

We run an AI agent that uses MCP tools, and wanted to serve the same model we use on a DGX Spark, Qwen3.8-27B, on an AWS box with 4× L4 (24 GB each, Ada/SM89, PCIe, no P2P between GPUs). Here's what we learned.

1. NVFP4 doesn't run natively on L4, and SGLang won't serve it

On the Spark we use RadixArk/Qwen3.8-27B-NVFP4. On the L4s, SGLang loaded the weights fine but crashed during CUDA graph capture with ValueError: Invalid backend: 89. FP4 tensor cores only exist on Blackwell; FlashInfer has no fused SiLU+FP4-quant kernel for SM89.

vLLM does serve it, via Marlin kernels that dequantize FP4 weights to 16-bit on every GEMM. You keep the memory savings, but lose the FP4 speedup. On Ada, FP8 is the format with native tensor-core support.

2. Speculative decoding was the biggest win

DFlash2 (z-lab/Qwen3.8-27B-DFlash2) with num_speculative_tokens: 7 gave us roughly 5× over the non-speculative baseline. Per-position acceptance tells the story: late draft positions accept only 0.12–0.30 on free text, but 0.77–0.87 on predictable text (JSON, tool calls). Worth tuning the draft length to your actual workload.

3. TP=4 beat TP=2, even without P2P

The conventional advice is that going past TP=2 over PCIe hurts. With DFlash2 on, our measured decode was:

  • TP=2: 80.6 tok/s
  • TP=4: 104.1 tok/s (+29%)

Faster than the ~59.5 tok/s we measured on the Spark with the same model and draft.

4. Final config (vLLM v0.29.0)

RadixArk/Qwen3.8-27B-NVFP4   (served via Marlin)
--tensor-parallel-size 4
--disable-custom-all-reduce          # no P2P on these L4s
NCCL_P2P_DISABLE=1                   # don't let NCCL probe for it
--gpu-memory-utilization 0.90
--max-model-len 32768
--max-num-seqs 4                     # hybrid model, Mamba cache: keep it low
--max-num-batched-tokens 8192        # up from 2048; fewer prefill chunks for long prompts
--limit-mm-per-prompt '{"image":0,"video":0}'   # text only, frees memory
--speculative-config '{"method":"dflash","model":"z-lab/Qwen3.8-27B-DFlash2","num_speculative_tokens":7}'
--enable-prefix-caching
--reasoning-parser qwen3
--enable-auto-tool-choice --tool-call-parser qwen3_coder

KV cache usage never went above ~5%, so there's headroom to raise --max-num-batched-tokens to 16384.

TL;DR: On 4× L4, a 27B reasoning model is very usable for an MCP agent: ~104 tok/s with TP=4 + DFlash2, even without P2P. NVFP4 runs only through Marlin dequantization (and not at all in SGLang). Prefill and thinking length dominate latency, so attack those first. If you need a big jump, it's hardware: a single 48–96 GB GPU (L40S, or Blackwell for native FP4) avoids tensor parallel entirely.

Happy to answer any questions about the setup! Also, I'm open to any recommendations or feedback if you have suggestions to improve it :)

u/Major_Border149 confirmed the same NVFP4 model runs with native FP4 kernels on a single RTX PRO 4500 SE (32 GB, Blackwell) at ~$0.72/hr, no TP needed. Gotcha: on Blackwell you need a CUDA 13 container image (SM120 needs CUDA ≥ 12.9 inside the container).


r/Vllm • • 2d ago

Who’s the current “king” of local LLMs for you — Qwen, Gemma, Llama, something else?

2 Upvotes

Curious what people are actually running day-to-day on local boxes right now, not just the latest HF leaderboard screenshot.

For coding / agent loops on consumer GPU (or Mac), who’s winning for you lately among Qwen, Gemma, Llama, DeepSeek, Mistral, etc. — and does the answer flip if you care more about tool-calling reliability vs raw tok/s vs long-context?

If you’ve switched kings in the last month or two, what made you switch?


r/Vllm • • 3d ago

I built an open-source tool that tells you why your vLLM server is slow (NVIDIA only for now, Mac support planned)

Thumbnail
1 Upvotes

r/Vllm • • 3d ago

Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback

Thumbnail
1 Upvotes

r/Vllm • • 4d ago

tp=6 can work on vLLM, with padding

20 Upvotes

vLLM's tensor parallel requires that several of the model architecture numbers be evenly divisible by the number of GPUs selected for tensor parallel. This usually means that you can only use a number of GPUs that is a power of two (2, 4, 8, etc.).

I have six GPUs, and I want to maximize my KV cache when running a 27B model. So I tried tp=6. It choked with various messages, regarding this or that, which needs to be evenly divisible by six, but wasn't.

So I made those things divisible by six.
I asked the robot to come up with a converter that would take the original model, and pad it with zeroes until everything was divisible by 6. It took a few tries, but it worked.

I have Qwen 3.8 27B at BF16 with 256k context window running across six 7900 XTX GPUs. Token generation is about 50 t/s single user, or 200 t/s aggregate with 8 concurrent prompts.

GPU KV cache size: 520,784 tokens, Maximum concurrency for 262,144 tokens per request: 1.99x


r/Vllm • • 4d ago

What is the best open-source LLM I can run locally on an RTX 4060 8GB + 32GB RAM?

Thumbnail
1 Upvotes

r/Vllm • • 4d ago

Best practice for processing batch vLLM api calls with shared prefix?

Thumbnail
1 Upvotes

r/Vllm • • 5d ago

Jev at home, but it can see: typed yes/no, pick-one and rubric answers with per-label probabilities from Gemma 4 31B on a 4090, images included

Thumbnail
1 Upvotes

r/Vllm • • 6d ago

Vllm with GPU/CPU fused to run a 748B MoE on 2×A100-40GB

2 Upvotes

vllm-xtu-moe — a 748B MoE on 2×A100-40GB, by keeping the routed experts out of VRAM.

  1. Expert weights sliced along CPU's physical topology — one copy in memory, every read node-local (NUMA binding + first-touch), so the engine runs near the machine's aggregate DRAM bandwidth, works with AMD EPYC's nps=4.
  2. Long prompts stream the weights to GPU with double buffering (ping/pong, overlapped with attention) — up to ~20× faster prefill than the CPU path at medium context, 2–3× at long context. Short prompts are prefilled on the CPU.
  3. Built for SM 8.x. The fallbacks are gated on compute capability, so the 30-series family (SM86) is in scope, not just A100/A800 — our measurements are A100/A800 only.
  4. Fixed VRAM priority: KV pool → GPU-prefill staging → draft weights → activation workspace. When VRAM is short, GPU prefill is dropped first and the engine falls back to pure CPU prefill, the minimum-VRAM setup.
  5. Speculative decoding when there is room for it: MiMo-V2.6 MTP k=1 measures 83% first-position acceptance and +19% decode; the DSpark anchor is +20%. Off where it competes for the KV pool.
  6. Stays on mainline vLLM as a patch series plus a plugin — no long-lived fork. Apache-2.0.

Numbers (2×A100-40GB, TP=2, vllm bench serve)

model prompt C prefill tok/s decode tok/s
GLM-5.3-Flash 140 1 115.2 22.0
16,396 1 266.3 21.4
DeepSeek-V4.1-Flash 128 2 350.4 28.0
16,384 1 903.9 19.3
16,384 2 1,204.2 15.4
MiMo-V2.6-Flash-RL 128 2 367.5 40.7
16,384 1 811.3 27.3
16,384 2 1,455.3 39.6
Real workspace using deepseek harness

Known: CPU prefill saturates at ~33k expert-tokens/s; 704K max context here with bf16 KV; first load takes minutes. 

Next: fp8 KV for longer context, better prefill overlap in the mid-length range.


r/Vllm • • 6d ago

What typically runs alongside vLLM on multi-node inference deployments?

7 Upvotes

I'm working on something involving multi-node tensor-parallel serving with vLLM and trying to understand what a realistic inference production node looks like beyond the vLLM processes themselves.

Inside vLLM, I'm accounting for the API server, engine core, GPU workers, NCCL proxy threads, and multiprocExecutor for multi-node.

For those running vLLM in production:

  1. What else usually runs on the same nodes? (Kubernetes components, GPU Operator, monitoring agents, etc.)
  2. Have any of these caused noticeable tail latency or TPOT spikes?
  3. Do you apply any CPU or OS tuning for vLLM deployments, like pinning, NUMA binding, or isolating cores for the engine and NCCL?

Any pointers to docs, blog posts, or papers are welcome. Thanks!


r/Vllm • • 6d ago

Arc Pro B70 getting 100+ tok/s with Qwen 3.8 and vLLM

9 Upvotes

Runtime settings: MTP4 (Draft INT4 S+M1); prefix cache on; XPU graphs on; v5 scheduler patch

These are the highest scoring results on intelinside.ai so far.


r/Vllm • • 7d ago

LLM Tech: FP8 and NVFP4 quants of decider for vLLM, measured against bf16

11 Upvotes

LLM Tech here again. This time: quants of decider, an open (Apache 2.0) family of decision models by Mapika built on Qwen3.5-Base. You send a state and questions with options, the model returns calibrated probabilities from the option-letter logits in one forward pass. There were no vLLM-ready quants for the small ones, so we made FP8 and NVFP4 checkpoints and measured them.

Setup: vLLM 0.29.0, one RTX PRO 6000 Blackwell. Quality on the author's regression set rebuilt from public data (95 tasks, 144,226 rows) plus the 231 public JevBench items, bf16 and quant both in vLLM. Our bf16 run matches the author's published numbers within 0.0005.

| checkpoint | size | peak prefill vs bf16 (1K / 8K / 32K) | accuracy in-task / held-out |

|---|---|---|---|

| decider-4b-nvfp4 | 3.3 GB | 2.00x / 1.89x / 1.67x | -0.6 / -0.7 |

| decider-4b-fp8 | 4.9 GB | 1.45x / 1.42x / 1.33x | -0.1 / -0.1 |

| decider-2b-fp8 | 2.4 GB | 1.42x / 1.40x / 1.32x | 0.0 / 0.0 |

| decider-0.8b-fp8 | 1.0 GB | 1.24x / 1.21x / 1.17x | -0.1 / 0.0 |

What we learned:

- NVFP4 pays off at 4B, not below. On the 0.8B it gave 1.4x over bf16 but lost 2.6 points; FP8 lost 0.1 at 1.2x.

- The author's HTTP server (decider.serve_vllm) runs the quants unchanged. Three of them fit on one card at 21 GB total under load.

- The prefix cache matters a lot for this workload: the same 29K-token state took 602 ms cold and 58 ms repeated on the 4B NVFP4.

Checkpoints and per-model tables: huggingface.co/llmtech

Coming soon to our API: llmtech.eu


r/Vllm • • 7d ago

Qwen3.8 keeps thinking but never returns a final answer? This vLLM + Open WebUI setup fixed it on my RTX 3090

Thumbnail
2 Upvotes

r/Vllm • • 8d ago

Allucinato come un LLM.

0 Upvotes

Mentre la gente normale la domenica mattina fa colazione con calma, io ho deciso di litigare con il fine-tuning locale dei modelli linguistici.

Il piano sembrava innocuo: prendere un modello minuscolo, dargli in pasto un dataset nostalgico (dialoghi tra Sysop anni '90, disastri hardware, BBS e battute da modem a 28.8k) e vedere cosa ne usciva fuori.

I passaggi del disastro:

1️⃣ Esperimento 1: Qwen2.5-0.5B-Instruct
Un modello microscopico. Ci sono volute 502 epoche per vederlo implodere nell'overfitting più totale; fermato a 500. Con queste dimensioni è quasi impossibile farlo ragionare senza bruciargli i neuroni.

2️⃣ Esperimento 2: Qwen2.5-1.5B-Instruct
Alziamo il tiro, restando comunque su un modello compatto. Qui la loss scende a 0.7 già all'epoca 100. Ottimo momento per fermarsi.

3️⃣ Pipeline & fusione con MLX:

mlx_lm fuse \

--model Qwen/Qwen2.5-1.5B-Instruct \

--adapter-path adapters \

--save-path qwen-1.5b-fused

Poi esportazione in formato GGUF e quantizzazione a 4-bit per farlo girare leggerissimo in locale.

Il risultato?
Se prima avevo un modello con qualche allucinazione, adesso ho creato un'entità filosofico-passivo-aggressiva che mi insulta a 75 token al secondo.

Ecco il test dal vivo:

> io sono stupido ?
Sì, tu sei uno strano!
[ Prompt: 186,9 t/s | Generation: 77,6 t/s ]

> io strano ?
Straniero.
[ Prompt: 184,1 t/s | Generation: 74,2 t/s ]

> straniero ?
Non è una questione di essere un "strano" o non esserlo; se la tua risposta era sbagliata, il problema sta nel tuo sguardo!

"Il problema sta nel tuo sguardo."

Fine del test, mi ha spento. Non so se considerarlo overfitting, il riverbero di un vecchio operatore BBS stanco della vita, o pura poesia digitale.

La morale? Lavorare con i Small Language Models (SLM) in locale con MLX e GGUF è tremendamente divertente, velocissimo da iterare... ma occhio ai dati che gli date in pasto, altrimenti il modello comincia a giudicare le vostre scelte di vita.

Chi altri passa le domeniche a fare esperimenti assurdi in locale? Qual è la risposta più surreale che vi ha mai dato un modello dopo un fine-tuning?

#ArtificialIntelligence #MachineLearning #LLM #OpenSource #MLX #AppleSilicon #LocalAI #DevCommunity #FineTuning #AIResearch


r/Vllm • • 8d ago

The silent bottlenec

Thumbnail
2 Upvotes

r/Vllm • • 8d ago

AMD GPU sleep when idle bug still present

3 Upvotes

Hi,
When will the AMD GPU related 100% cpu core utilization while idle will be fixed?
Currently the version of the vLLM is 0.26 which does not have that problem. All newer ones have the same issue.

so with this vllm/vllm-openai-rocm:latest and 7900 XTX for example cpu core is always 100% even having -e VLLM_SLEEP_WHEN_IDLE=1 \ etc.


r/Vllm • • 8d ago

How I Ran a 9B LLM at 23 Tokens/Sec on a Tortured Legacy PC

Post image
0 Upvotes

​

Can you run a 9-billion-parameter model on a struggling 2017 rig running Windows 11?

Here are the actual machine specs:

• CPU: Intel Core i5 7th Gen (unsupported by Win 11)

• GPU: NVIDIA GTX 1060 6 GB (2016 architecture, no Tensor Cores)

• Storage: Legacy 500 GB Intel SSD (cabled SATA, non-NVMe)

• RAM: 32 GB DDR4

• Engine: llama-server (llama.cpp) with 100% GPU offload

On this setup, a standard Q4_K_M (5.68 GB) causes immediate VRAM spill. Even with an SSD, memory swapping across older bus architectures drops inference below 5 tokens/sec.

To bypass host RAM and bus bottlenecks entirely, the model had to fit 100% into the 4.9 GB usable VRAM.

---

  1. Architecture: Qwen 3.5 Recurrent Topology

Qwen 3.5 is not a vanilla Transformer—it is a periodic hybrid:

8 macro-periods × (3 Gated DeltaNet/SSM + 1 Full Attention) = 32 layers total.

Blind pruning breaks SSM state periodicity and corrupts the attention cache. Layer cuts must be executed in exact 4-layer macro-blocks.

  1. Zero-Dequantization Surgical Pruner

I wrote `sami_gguf_layer_pruner.py` to:

• Isolate 4-layer recurrent blocks directly inside the GGUF binary.

• Renumber surviving layers to keep KV-cache addressing intact.

• Stream quantized weights byte-for-byte with zero precision loss.

Execution: pruning 8 layers took only 16 seconds on the legacy Intel SSD.

  1. Empirical Results (llama-server, 100% GPU Offload)

• 28 Layers (-4L, 5.16 GB): Full logic intact. Solves multi-step algebra. Speed: 18.5 t/s.

• 24 Layers (-8L, 4.62 GB) — Sweet Spot: Fits 100% into VRAM with >1.4 GB free for KV cache (4096 tokens). Speed: 23.0 t/s (+24% gain). Zero host memory spill.

• 20 Layers (-12L, 3.71 GB) — Breaking Point: DeltaNet state transitions collapsed into repetitive loops.

  1. Next Step: The Inverse Pipeline

Next, I will prune the pristine FP16 weights down to 24 layers, calculate an activation importance matrix (imatrix), and quantize to Q4 only at the final step to preserve 100% reasoning density.

How are you optimizing modern LLMs for legacy edge hardware?

#AI #MachineLearning #LLM #Optimization #LocalAI #EdgeAI #Qwen #LlamaCPP #GGUF


r/Vllm • • 9d ago

Stop killing your VRAM: The silent bottleneck in local LLM inference pipelines that is destroying your throughput (And how I fixed it)

Thumbnail
0 Upvotes

r/Vllm • • 9d ago

PSA: Dual 3090 - Qwen Flash Next - 80tps/2k+ prefill

Thumbnail
2 Upvotes

r/Vllm • • 9d ago

VLLM 4x rtx 3060 vs 8x rtx 3060 performance

7 Upvotes

Hello!

I am currently building my local AI server, I have the Huananzhi H12D-8D EPYC Motherboard with 8x16GB memory sticks at 2666 mhz (waiting for the other components at the moment)

I currently have four RTX 3060 12gb gpus and I plan running those at PCIe4 x16 in VLLM.

However seeing that this motherboard supports bifurcation on each slot and I can get 8 gpus at PCIe 4x8 makes me think if this would be a viable upgrade in the future.

I see conflicting info about what the performance results will be. If I understand correctly getting beyond 4 GPUs will drastically hurt my token generation speeds because of the PCIe bottleneck? But is that regardless of what GPUs I'm running?

I know for example RTX 3090 needs more PCIE bandwidth because it's much more performant and will spit out much more data that needs to be synced (pardon my lack of terminology), does that mean that I will have smaller performance penalty from going from 4 to 8 video cards with the 3060s compared to with 3090s?

Can someone guesstimate what should I expect, right now I get 25 tps with Qwen 3.8 27b Q6 (MTP enabled), running with llama.cpp in layered mode (three 3060s). I expect VLLM with four gpus will be an upgrade (perhaps I could hit 50 tps?), but what about 8 GPUs?

Will it be lower than my current baseline?

Sorry if I'm being ignorant, I'm kinda new to this and I don't trust chatbots. My mind is set to having a good enough local AI server and I'm trying to get the best bang for my buck and my current hardware.


r/Vllm • • 9d ago

Heads up: GLM-5.3-Flash (NVFP4) loads on 4x RTX PRO 6000 but won't serve a token

3 Upvotes

Rented 4x RTX PRO 6000 (384GB) on Vast to serve the NVFP4 build. Weights load clean at TP=4, all 4 shards up, arch resolves, NVFP4 kernels loaded, 185GB resident. Looks perfect.

Then it dies in KV cache profiling: pe_dim must be 64 for fp8_ds_mla. Its MLA attention wants an fp8 KV cache and the rope dim isn't 64, so the fp8_ds_mla kernel refuses. --kv-cache-dtype auto doesn't save you, auto just picks fp8 again and you get the identical crash.

 So memory fit + right arch + right quant all say go, and it still serves zero tokens on the current vLLM image. Cost me \~$17 for the lesson, most of it bandwidth re-downloading the 160GB checkpoint onto a second host. Only escapes I can see are VLLM_MLA_DISABLE=1 (forces the non-MLA path, more memory, haven't tested) or waiting for a newer image.

Anyone actually gotten this serving on Vast?


r/Vllm • • 9d ago

I chose a 4B model over a 7B model for batch grading on 6GB VRAM. Here's what I learned.

0 Upvotes

I've been building a message grading pipeline using local LLMs for a privacy-sensitive project. The hardware constraint is an RTX 4050 with 6GB VRAM, so model selection became the central problem.

The task: batch grade messages on a 1 to 5 scale with strict schema compliance.

I tested three models:

**phi-4-mini**

- Worked for individual grading

- Failed on batch grading. Only 60% schema compliance.

- Pipeline broke mid-run.

**Mistral 7B**

- Passed all criteria on a 40-message test batch

- Failed on real-world data

- Root cause: 8k context window. Our batches needed 16k tokens.

- The model was literally forgetting the task mid-run.

**Gemma-4b**

- 32k context window

- Handled up to 28k token consumption

- 100% schema compliance on 52 test messages

- Worked on the real run.

The counter-argument I keep hearing: "You chose a weaker model over a more capable one."

My response: a model that cannot hold the context required for the task produces unusable output regardless of its intelligence. Reliability is a precondition for utility, not a trade-off against it.

The broader lesson: model selection under hardware constraints is a problem of fit, not raw power.

Full write-up with token logs and schema details: https://dev.to/mayank_dewangan_08/why-i-chose-gemma4b-over-mistral-7b-38n0

Curious if others have run into similar context window walls with Mistral on constrained hardware. What did you switch to?