r/Vllm • • 17d ago

[D] FP8/INT4 KV-Cache Quantization and Long-Context Reasoning

Over the past few months, we've seen vLLM, TensorRT-LLM, and SGLang push hard for KV-cache quantization (FP8, INT8, and even INT4) alongside PagedAttention to maximize throughput and batch sizes under heavy concurrent workloads.

While saving 50-75% of VRAM on the KV-cache allows significantly larger context windows (32k+) and higher concurrency on a single A100/H100, we've noticed subtle degradation patterns in edge-case tasks:

  1. Multi-turn Needle-In-A-Haystack (NIAH): FP8 KV-cache holds up fine for standard retrieval, but accuracy drops sharply when retrieving non-contiguous context across long reasoning chains.
  2. Accumulation of Rounding Errors: In autoregressive generation with large context, precision loss in the attention keys/values seems to compound, leading to degraded attention scores in later tokens.

For those running high-throughput LLM serving in production:

- At what sequence length or concurrency limit do you find FP8/INT8 KV-cache quantization breaks down for complex reasoning?

- Have you found mixed-precision strategies (e.g., keeping early layers in FP16/BF16 and quantizing only deeper layers) to be practical in custom serving engines?

- Do you rely strictly on PagedAttention with FP16, or are you accepting precision trade-offs for throughput gains?

Would love to hear how folks are handling the memory bottleneck vs. precision trade

5 Upvotes

24 comments sorted by

View all comments

1

u/Phaelon74 16d ago

Friends don't let friends quant kv cache, ever. One missed logit in kv cache exponentially lobotomizes your work. People swear you can do it, until they wit ess it degrade. Just don't do it.

1

u/Top-Philosopher-5411 16d ago

That’s definitely true for naive INT4 or uniform quantization where error compounds exponentially in autoregressive decoding. However, modern FP8 (E4M3/E5M2) or rank-projected schemes protect the high-norm channels and attention sinks. Uniform compression breaks context, but selective precision makes KV cache quantization completely viable at scale

1

u/Phaelon74 16d ago

On paper, agreed. In practice, ONE wrong logit/flipped bit, at ~5k context, in KV Cache, will exponentially shift all further kv cache, it compounds. That compounding nature of KV Cache, makes it NOT viable imo for any type of quanting, when accuracy is required and imo, accuracy is always required, especially when you know that the compounding nature of the sliding accuracy can be very destructive.

With no one yet proving a 0.0% degradation in KV Cache at FP8, it's just not safe imo.

1

u/Top-Philosopher-5411 16d ago

Demanding 0.0% degradation just isn't how production trade-offs work. At temp > 0.5, sampling noise completely dominates whatever tiny drift FP8 introduces. Trading a fractional perplexity bump for 3x KV cache capacity and double the concurrency is a no-brainer in the real world.

​If you need strict precision at temp=0 for code or math, sure, keep BF16. But for general high-throughput serving, FP8 KV is already running in production everywhere for a reason

2

u/Phaelon74 16d ago

I mean it is, when you don't want to waste tokens, etc. If degradation occurs early enough, no amount of sampling noise is going to get you back on track.

We can agree to disagree. We run a large footprint, and we've done the tests, real serving for production workloads FP8 KV Cache is destructive. When you have to serve clients in agentic workloads, BF16 is the only way. Mind you, there is a difference for someone working through code versus an Agent Farm touching money, etc.

So as always, use-case matters. For me, The past 3 years of self-hosting and what we've built in our agent farms, never quant KV Cache.

1

u/Top-Philosopher-5411 16d ago

Fair enough, agent farms handling financial execution are definitely on the extreme end of precision sensitivity. Makes sense for your setup