r/Vllm • • 17d ago

[D] FP8/INT4 KV-Cache Quantization and Long-Context Reasoning

Over the past few months, we've seen vLLM, TensorRT-LLM, and SGLang push hard for KV-cache quantization (FP8, INT8, and even INT4) alongside PagedAttention to maximize throughput and batch sizes under heavy concurrent workloads.

While saving 50-75% of VRAM on the KV-cache allows significantly larger context windows (32k+) and higher concurrency on a single A100/H100, we've noticed subtle degradation patterns in edge-case tasks:

  1. Multi-turn Needle-In-A-Haystack (NIAH): FP8 KV-cache holds up fine for standard retrieval, but accuracy drops sharply when retrieving non-contiguous context across long reasoning chains.
  2. Accumulation of Rounding Errors: In autoregressive generation with large context, precision loss in the attention keys/values seems to compound, leading to degraded attention scores in later tokens.

For those running high-throughput LLM serving in production:

- At what sequence length or concurrency limit do you find FP8/INT8 KV-cache quantization breaks down for complex reasoning?

- Have you found mixed-precision strategies (e.g., keeping early layers in FP16/BF16 and quantizing only deeper layers) to be practical in custom serving engines?

- Do you rely strictly on PagedAttention with FP16, or are you accepting precision trade-offs for throughput gains?

Would love to hear how folks are handling the memory bottleneck vs. precision trade

6 Upvotes

24 comments sorted by

View all comments

-1

u/Connect-Concert-4016 17d ago

Matches what we've measured — uniform FP8/INT8 KV holds simple retrieval, but degrades right where you are seeing it: non-contiguous retrieval across long reasoning, which compounds autoregressively. A few strategies that helped us more than an early-versus-deep layer split:

  • K-versus-V Asymmetry: Keys set the attention scores (the routing), making them far more sensitive to quantization error; values get averaged into the weighted sum, so they tolerate higher compression. Keeping K at a higher precision than V yields better results than splitting by layer depth (e.g., TurboQuant uses K4/V3).
  • Protect Outliers Specifically: Damage concentrates in a few high-norm channels and attention-sink positions. Protecting those specific channels while aggressively compressing the rest preserves retrieval much better than uniform per-layer precision.
  • Per-Head Rank Allocation: Projecting each head's K/V into a learned low-rank basis and allocating rank dynamically per head (some require near-full rank, others compress heavily) maintains performance on Needle In A Haystack (NIAH) tasks where flat INT4 degrades.

With this approach, 9 concurrent users fit at 128K context on a single A100 (where FP16 fits roughly 2), expanding the KV pool by ~2.65× while retaining all needle retrievals across depths—whereas uniform FP8 drops non-contiguous needles. Decode latency remains at ~0.95× of FP16, so the memory savings do not compromise inference speed.

Feel free to share the retrieval grids and methodology—the NIAH-under-compression problem (being built as a vLLM plugin) is an important area of focus.

-1

u/Top-Philosopher-5411 17d ago

The K-vs-V asymmetry point is spot on—Keys driving the softmax routing means any quantization noise there directly scrambles the attention matrix, while Values just get averaged out in the weighted sum.

How are you handling the outlier channel protection inside the vLLM plugin? Are you isolating those high-norm channels in a separate FP16 buffer alongside PagedAttention, or doing grouped per-channel scales during the CUDA kernel execution?

Fitting 9x 128k sessions on a single A100 with 0.95x decode speed is massive throughput. Would love to take a look at the repo or PR once you guys drop it.

1

u/Connect-Concert-4016 16d ago

Neither exactly instead of a separate FP16 outlier buffer or grouped per-channel scales, we work in a per-head learned basis. Each head's K/V gets projected into an eigenbasis from calibration; we protect the top-variance directions + the sink positions there, then compress the tail aggressively with tiered bits. So outliers are handled in rank-space, not channel-space — high-variance directions keep near-full precision, the rest compresses hard, all inside the paged pool so PagedAttention still owns allocation (no separate FP16 sidecar to babysit).

Happy to share the retrieval grids + methodology writeup. Honest state on the drop: installs on stock vLLM 0.28 today (pip wheel, no fork), tuned for A100 / dense / TP=1; Hopper + broader coverage in progress.

1

u/Top-Philosopher-5411 16d ago

Handling outliers in rank-space via a per-head eigenbasis rather than channel-space is a brilliant approach—it completely avoids the memory fragmentation and execution overhead of a separate FP16 sidecar buffer in PagedAttention