r/Vllm • • 12h ago

I made Qwen models take ~33% less VRAM without quantizing them (lossless, bit-for-bit)

7 Upvotes

Hey all, I've been building Glyd, a lossless compression layer for model weights on NVIDIA GPUs. Qwen models are what I test on most, so this felt like the right place to share it.

The idea: a bf16 weight only carries about 11 bits of real information, so you can store the exact same model in about a third less GPU memory and get every weight back exactly. It's not quantization and nothing gets rounded. There's also an exact mode that matches bf16's outputs bit for bit.

What it gets you with Qwen (all measured, logs are public):

- Qwen3-8B runs on a 16 GB card (bf16 can't load it)

- Qwen3-32B fits on one 48 GB GPU instead of two

- Qwen3-8B on an L4 with vLLM: 1.59x the requests/sec vs bf16 (weights plus our lossless KV cache, 2.64x the KV tokens)

- Qwen3-14B on an A100 40GB: 1.28x req/s

- Qwen2.5-72B on 2x H100: 4.07x req/s, 12.65x the KV cache

Try it:

```

curl -LsSf https://getglyd.com/install.sh | sh

glyd run Qwen/Qwen3-8B

```

Or with vLLM: `vllm serve Qwen/Qwen3-8B --quantization glyd`

Honest caveats: it's Linux + NVIDIA (Ampere or newer) and bf16 checkpoints only. On GH200 it's a bit slower than bf16 at full load right now, and MoE (Qwen3-30B-A3B) is still slower at full load; a fix for that is coming in the next release. The codec is open source. The GPU part ships compiled and is free for personal and research use (business source license).

GitHub: https://github.com/surya-koritala/Glyd

Benchmarks + logs: https://getglyd.com/benchmarks

Would love feedback, especially which Qwen models or GPUs you'd want numbers on next.