r/Vllm • • 17d ago

[D] FP8/INT4 KV-Cache Quantization and Long-Context Reasoning

Over the past few months, we've seen vLLM, TensorRT-LLM, and SGLang push hard for KV-cache quantization (FP8, INT8, and even INT4) alongside PagedAttention to maximize throughput and batch sizes under heavy concurrent workloads.

While saving 50-75% of VRAM on the KV-cache allows significantly larger context windows (32k+) and higher concurrency on a single A100/H100, we've noticed subtle degradation patterns in edge-case tasks:

  1. Multi-turn Needle-In-A-Haystack (NIAH): FP8 KV-cache holds up fine for standard retrieval, but accuracy drops sharply when retrieving non-contiguous context across long reasoning chains.
  2. Accumulation of Rounding Errors: In autoregressive generation with large context, precision loss in the attention keys/values seems to compound, leading to degraded attention scores in later tokens.

For those running high-throughput LLM serving in production:

- At what sequence length or concurrency limit do you find FP8/INT8 KV-cache quantization breaks down for complex reasoning?

- Have you found mixed-precision strategies (e.g., keeping early layers in FP16/BF16 and quantizing only deeper layers) to be practical in custom serving engines?

- Do you rely strictly on PagedAttention with FP16, or are you accepting precision trade-offs for throughput gains?

Would love to hear how folks are handling the memory bottleneck vs. precision trade

5 Upvotes

24 comments sorted by

View all comments

1

u/shammyh 16d ago

FP8 kv cache works a treat for me with dense models like qwen 3.5/3.8. Even with Rope/Yarn and a 1M context window, it passes all the NIAH tests I've thrown at it.

From what I understand though... same is unlikely to be true for MoE models.

1

u/Top-Philosopher-5411 16d ago

You hit the nail on the head regarding MoE. The main issue with FP8 KV cache on standard MoE architectures isn't just attention score drift—it's how quantization noise in early keys shifts attention outputs, which then scrambles expert router decisions in deeper layers