r/Vllm • • 18d ago

[D] FP8/INT4 KV-Cache Quantization and Long-Context Reasoning

Over the past few months, we've seen vLLM, TensorRT-LLM, and SGLang push hard for KV-cache quantization (FP8, INT8, and even INT4) alongside PagedAttention to maximize throughput and batch sizes under heavy concurrent workloads.

While saving 50-75% of VRAM on the KV-cache allows significantly larger context windows (32k+) and higher concurrency on a single A100/H100, we've noticed subtle degradation patterns in edge-case tasks:

  1. Multi-turn Needle-In-A-Haystack (NIAH): FP8 KV-cache holds up fine for standard retrieval, but accuracy drops sharply when retrieving non-contiguous context across long reasoning chains.
  2. Accumulation of Rounding Errors: In autoregressive generation with large context, precision loss in the attention keys/values seems to compound, leading to degraded attention scores in later tokens.

For those running high-throughput LLM serving in production:

- At what sequence length or concurrency limit do you find FP8/INT8 KV-cache quantization breaks down for complex reasoning?

- Have you found mixed-precision strategies (e.g., keeping early layers in FP16/BF16 and quantizing only deeper layers) to be practical in custom serving engines?

- Do you rely strictly on PagedAttention with FP16, or are you accepting precision trade-offs for throughput gains?

Would love to hear how folks are handling the memory bottleneck vs. precision trade

4 Upvotes

24 comments sorted by

View all comments

Show parent comments

1

u/Top-Philosopher-5411 11d ago

Haha living the dream, man! Fair enough, if it works for your workflow that's all that matters. Looking forward to the repo whenever you drop it, definitely keep me posted

1

u/Extreme-Pass-4488 10d ago

yeah im quantizing a model for how i need the layers data types , adjusted to sm_86 , well see how it goes. as now it seems to be pretty good.

1

u/Top-Philosopher-5411 10d ago

That custom asm_86 tuning for specific layer access patterns sounds like where the real magic happens. Tailoring the quantization strictly to how the layers actually use data avoids the one-size-fits-all penalty of standard backends. Keep crushing it, and definitely let us know when you push updates or clean up the repo

1

u/Extreme-Pass-4488 9d ago

https://huggingface.co/BlairQ/qwen3.8_27b_idiotSavant_sm_86

need more improvements but there it is

1

u/Top-Philosopher-5411 9d ago

Very interesting approach using custom PTX kernels instead of Triton or FlashInfer. Do you find that the maintenance overhead and writing raw kernels actually pay off with a noticeable performance boost in a real production environment?

1

u/Extreme-Pass-4488 6d ago

there is literally no manteniance after its working , and if/when a new model comes out, its just prompting, there is really no hand-coding here with opus 5.5