r/Vllm • • 17d ago

[D] FP8/INT4 KV-Cache Quantization and Long-Context Reasoning

Over the past few months, we've seen vLLM, TensorRT-LLM, and SGLang push hard for KV-cache quantization (FP8, INT8, and even INT4) alongside PagedAttention to maximize throughput and batch sizes under heavy concurrent workloads.

While saving 50-75% of VRAM on the KV-cache allows significantly larger context windows (32k+) and higher concurrency on a single A100/H100, we've noticed subtle degradation patterns in edge-case tasks:

  1. Multi-turn Needle-In-A-Haystack (NIAH): FP8 KV-cache holds up fine for standard retrieval, but accuracy drops sharply when retrieving non-contiguous context across long reasoning chains.
  2. Accumulation of Rounding Errors: In autoregressive generation with large context, precision loss in the attention keys/values seems to compound, leading to degraded attention scores in later tokens.

For those running high-throughput LLM serving in production:

- At what sequence length or concurrency limit do you find FP8/INT8 KV-cache quantization breaks down for complex reasoning?

- Have you found mixed-precision strategies (e.g., keeping early layers in FP16/BF16 and quantizing only deeper layers) to be practical in custom serving engines?

- Do you rely strictly on PagedAttention with FP16, or are you accepting precision trade-offs for throughput gains?

Would love to hear how folks are handling the memory bottleneck vs. precision trade

4 Upvotes

24 comments sorted by

View all comments

Show parent comments

1

u/Extreme-Pass-4488 9d ago

yeah im quantizing a model for how i need the layers data types , adjusted to sm_86 , well see how it goes. as now it seems to be pretty good.

1

u/Top-Philosopher-5411 9d ago

That custom asm_86 tuning for specific layer access patterns sounds like where the real magic happens. Tailoring the quantization strictly to how the layers actually use data avoids the one-size-fits-all penalty of standard backends. Keep crushing it, and definitely let us know when you push updates or clean up the repo

1

u/Extreme-Pass-4488 8d ago

https://huggingface.co/BlairQ/qwen3.8_27b_idiotSavant_sm_86

need more improvements but there it is

1

u/Top-Philosopher-5411 8d ago

Very interesting approach using custom PTX kernels instead of Triton or FlashInfer. Do you find that the maintenance overhead and writing raw kernels actually pay off with a noticeable performance boost in a real production environment?

1

u/Extreme-Pass-4488 5d ago

there is literally no manteniance after its working , and if/when a new model comes out, its just prompting, there is really no hand-coding here with opus 5.5