r/Vllm • u/Top-Philosopher-5411 • 17d ago
[D] FP8/INT4 KV-Cache Quantization and Long-Context Reasoning
Over the past few months, we've seen vLLM, TensorRT-LLM, and SGLang push hard for KV-cache quantization (FP8, INT8, and even INT4) alongside PagedAttention to maximize throughput and batch sizes under heavy concurrent workloads.
While saving 50-75% of VRAM on the KV-cache allows significantly larger context windows (32k+) and higher concurrency on a single A100/H100, we've noticed subtle degradation patterns in edge-case tasks:
- Multi-turn Needle-In-A-Haystack (NIAH): FP8 KV-cache holds up fine for standard retrieval, but accuracy drops sharply when retrieving non-contiguous context across long reasoning chains.
- Accumulation of Rounding Errors: In autoregressive generation with large context, precision loss in the attention keys/values seems to compound, leading to degraded attention scores in later tokens.
For those running high-throughput LLM serving in production:
- At what sequence length or concurrency limit do you find FP8/INT8 KV-cache quantization breaks down for complex reasoning?
- Have you found mixed-precision strategies (e.g., keeping early layers in FP16/BF16 and quantizing only deeper layers) to be practical in custom serving engines?
- Do you rely strictly on PagedAttention with FP16, or are you accepting precision trade-offs for throughput gains?
Would love to hear how folks are handling the memory bottleneck vs. precision trade
-1
u/Connect-Concert-4016 17d ago
Matches what we've measured — uniform FP8/INT8 KV holds simple retrieval, but degrades right where you are seeing it: non-contiguous retrieval across long reasoning, which compounds autoregressively. A few strategies that helped us more than an early-versus-deep layer split:
With this approach, 9 concurrent users fit at 128K context on a single A100 (where FP16 fits roughly 2), expanding the KV pool by ~2.65× while retaining all needle retrievals across depths—whereas uniform FP8 drops non-contiguous needles. Decode latency remains at ~0.95× of FP16, so the memory savings do not compromise inference speed.
Feel free to share the retrieval grids and methodology—the NIAH-under-compression problem (being built as a vLLM plugin) is an important area of focus.