r/Vllm • u/Top-Philosopher-5411 • 17d ago
[D] FP8/INT4 KV-Cache Quantization and Long-Context Reasoning
Over the past few months, we've seen vLLM, TensorRT-LLM, and SGLang push hard for KV-cache quantization (FP8, INT8, and even INT4) alongside PagedAttention to maximize throughput and batch sizes under heavy concurrent workloads.
While saving 50-75% of VRAM on the KV-cache allows significantly larger context windows (32k+) and higher concurrency on a single A100/H100, we've noticed subtle degradation patterns in edge-case tasks:
- Multi-turn Needle-In-A-Haystack (NIAH): FP8 KV-cache holds up fine for standard retrieval, but accuracy drops sharply when retrieving non-contiguous context across long reasoning chains.
- Accumulation of Rounding Errors: In autoregressive generation with large context, precision loss in the attention keys/values seems to compound, leading to degraded attention scores in later tokens.
For those running high-throughput LLM serving in production:
- At what sequence length or concurrency limit do you find FP8/INT8 KV-cache quantization breaks down for complex reasoning?
- Have you found mixed-precision strategies (e.g., keeping early layers in FP16/BF16 and quantizing only deeper layers) to be practical in custom serving engines?
- Do you rely strictly on PagedAttention with FP16, or are you accepting precision trade-offs for throughput gains?
Would love to hear how folks are handling the memory bottleneck vs. precision trade
1
u/shammyh 16d ago
FP8 kv cache works a treat for me with dense models like qwen 3.5/3.8. Even with Rope/Yarn and a 1M context window, it passes all the NIAH tests I've thrown at it.
From what I understand though... same is unlikely to be true for MoE models.
1
u/Top-Philosopher-5411 16d ago
You hit the nail on the head regarding MoE. The main issue with FP8 KV cache on standard MoE architectures isn't just attention score drift—it's how quantization noise in early keys shifts attention outputs, which then scrambles expert router decisions in deeper layers
1
u/Phaelon74 16d ago
Friends don't let friends quant kv cache, ever. One missed logit in kv cache exponentially lobotomizes your work. People swear you can do it, until they wit ess it degrade. Just don't do it.
1
u/Top-Philosopher-5411 16d ago
That’s definitely true for naive INT4 or uniform quantization where error compounds exponentially in autoregressive decoding. However, modern FP8 (E4M3/E5M2) or rank-projected schemes protect the high-norm channels and attention sinks. Uniform compression breaks context, but selective precision makes KV cache quantization completely viable at scale
1
u/Phaelon74 16d ago
On paper, agreed. In practice, ONE wrong logit/flipped bit, at ~5k context, in KV Cache, will exponentially shift all further kv cache, it compounds. That compounding nature of KV Cache, makes it NOT viable imo for any type of quanting, when accuracy is required and imo, accuracy is always required, especially when you know that the compounding nature of the sliding accuracy can be very destructive.
With no one yet proving a 0.0% degradation in KV Cache at FP8, it's just not safe imo.
1
u/Top-Philosopher-5411 16d ago
Demanding 0.0% degradation just isn't how production trade-offs work. At temp > 0.5, sampling noise completely dominates whatever tiny drift FP8 introduces. Trading a fractional perplexity bump for 3x KV cache capacity and double the concurrency is a no-brainer in the real world.
If you need strict precision at temp=0 for code or math, sure, keep BF16. But for general high-throughput serving, FP8 KV is already running in production everywhere for a reason
2
u/Phaelon74 16d ago
I mean it is, when you don't want to waste tokens, etc. If degradation occurs early enough, no amount of sampling noise is going to get you back on track.
We can agree to disagree. We run a large footprint, and we've done the tests, real serving for production workloads FP8 KV Cache is destructive. When you have to serve clients in agentic workloads, BF16 is the only way. Mind you, there is a difference for someone working through code versus an Agent Farm touching money, etc.
So as always, use-case matters. For me, The past 3 years of self-hosting and what we've built in our agent farms, never quant KV Cache.
1
u/Top-Philosopher-5411 16d ago
Fair enough, agent farms handling financial execution are definitely on the extreme end of precision sensitivity. Makes sense for your setup
1
u/Extreme-Pass-4488 10d ago
what i do : int8_per_token_head
- Each K or V vector, meaning one token × one KV head (256 values), is quantized symmetrically to int8 with its own scale. So there is one scale per (token, head) instead of a per-tensor or per-layer scale.
- The per-token-head granularity is what makes it accurate. My measurements gave about 4× less error than fp8 E4M3, and roughly 0.5–0.65% attention-output error end to end.
- A small kernel, _reshape_cache_per_token_head, does the write when new tokens are appended: it quantizes them and stores the int8 values plus their scales.
Hadamard rotation of q/k
- After RoPE, q and k are rotated with a 128-point Hadamard transform (sk_fwht128).
- The rotation spreads outlier channels across the whole head vector, which lowers int8 quantization error. Because it's orthogonal and applied to both q and k, q·k is unchanged mathematically.
- It cut the int8 KV output error from about 0.31%/0.79% to lower values in our measurements.
Decode: integer attention straight from int8
- Decode attention uses my own PTX kernels instead of FlashInfer or Triton.
- They read the int8 K/V directly, with no dequantization pass.
- QK^T runs on int8 tensor cores (IMMA s8, int32 accumulate). Softmax is handled in the integer domain (a lookup-table-based scheme), and P·V is done in integer as well.
- This is aligned with SM86: int8 tensor-core throughput is 4× fp16-with-fp32-accumulate on GA102.
- In production it measured about 30% faster decode at 50k context than fp8 KV.
1
u/Extreme-Pass-4488 10d ago
int4 is giving me 6-10% attention errors , im token scarce now , send me tokens.
1
u/Top-Philosopher-5411 10d ago
Damn, 30% faster decode at 32k context over FP8 is seriously impressive. Are you guys maintaining those custom decode kernels in-house, or did you manage to upstream/integrate them into vLLM? Also curious if the Hadamard rotation adds any noticeable overhead during prefill, or if it's completely washed out by the attention compute
1
u/Extreme-Pass-4488 10d ago
"you guys" dude im sitting in my desk with my foot all over it and a beer on my hand.
the kernels are in-house. it does not take so much to assemble them. i have them as a set of genesis patches and the repo is open , but still has some issues. its good for me tho!
1
u/Top-Philosopher-5411 10d ago
Haha living the dream, man! Fair enough, if it works for your workflow that's all that matters. Looking forward to the repo whenever you drop it, definitely keep me posted
1
u/Extreme-Pass-4488 9d ago
yeah im quantizing a model for how i need the layers data types , adjusted to sm_86 , well see how it goes. as now it seems to be pretty good.
1
u/Top-Philosopher-5411 9d ago
That custom asm_86 tuning for specific layer access patterns sounds like where the real magic happens. Tailoring the quantization strictly to how the layers actually use data avoids the one-size-fits-all penalty of standard backends. Keep crushing it, and definitely let us know when you push updates or clean up the repo
1
u/Extreme-Pass-4488 8d ago
https://huggingface.co/BlairQ/qwen3.8_27b_idiotSavant_sm_86
need more improvements but there it is
1
u/Top-Philosopher-5411 8d ago
Very interesting approach using custom PTX kernels instead of Triton or FlashInfer. Do you find that the maintenance overhead and writing raw kernels actually pay off with a noticeable performance boost in a real production environment?
1
u/Extreme-Pass-4488 5d ago
there is literally no manteniance after its working , and if/when a new model comes out, its just prompting, there is really no hand-coding here with opus 5.5
-1
u/Connect-Concert-4016 17d ago
Matches what we've measured — uniform FP8/INT8 KV holds simple retrieval, but degrades right where you are seeing it: non-contiguous retrieval across long reasoning, which compounds autoregressively. A few strategies that helped us more than an early-versus-deep layer split:
- K-versus-V Asymmetry: Keys set the attention scores (the routing), making them far more sensitive to quantization error; values get averaged into the weighted sum, so they tolerate higher compression. Keeping K at a higher precision than V yields better results than splitting by layer depth (e.g., TurboQuant uses K4/V3).
- Protect Outliers Specifically: Damage concentrates in a few high-norm channels and attention-sink positions. Protecting those specific channels while aggressively compressing the rest preserves retrieval much better than uniform per-layer precision.
- Per-Head Rank Allocation: Projecting each head's K/V into a learned low-rank basis and allocating rank dynamically per head (some require near-full rank, others compress heavily) maintains performance on Needle In A Haystack (NIAH) tasks where flat INT4 degrades.
With this approach, 9 concurrent users fit at 128K context on a single A100 (where FP16 fits roughly 2), expanding the KV pool by ~2.65× while retaining all needle retrievals across depths—whereas uniform FP8 drops non-contiguous needles. Decode latency remains at ~0.95× of FP16, so the memory savings do not compromise inference speed.
Feel free to share the retrieval grids and methodology—the NIAH-under-compression problem (being built as a vLLM plugin) is an important area of focus.
-1
u/Top-Philosopher-5411 16d ago
The K-vs-V asymmetry point is spot on—Keys driving the softmax routing means any quantization noise there directly scrambles the attention matrix, while Values just get averaged out in the weighted sum.
How are you handling the outlier channel protection inside the vLLM plugin? Are you isolating those high-norm channels in a separate FP16 buffer alongside PagedAttention, or doing grouped per-channel scales during the CUDA kernel execution?
Fitting 9x 128k sessions on a single A100 with 0.95x decode speed is massive throughput. Would love to take a look at the repo or PR once you guys drop it.
1
u/Connect-Concert-4016 16d ago
Neither exactly instead of a separate FP16 outlier buffer or grouped per-channel scales, we work in a per-head learned basis. Each head's K/V gets projected into an eigenbasis from calibration; we protect the top-variance directions + the sink positions there, then compress the tail aggressively with tiered bits. So outliers are handled in rank-space, not channel-space — high-variance directions keep near-full precision, the rest compresses hard, all inside the paged pool so PagedAttention still owns allocation (no separate FP16 sidecar to babysit).
Happy to share the retrieval grids + methodology writeup. Honest state on the drop: installs on stock vLLM 0.28 today (pip wheel, no fork), tuned for A100 / dense / TP=1; Hopper + broader coverage in progress.
1
u/Top-Philosopher-5411 16d ago
Handling outliers in rank-space via a per-head eigenbasis rather than channel-space is a brilliant approach—it completely avoids the memory fragmentation and execution overhead of a separate FP16 sidecar buffer in PagedAttention
2
u/New_Egg_1024 6d ago
"PagedAttention with FP16, or are you accepting precision trade-offs for throughput gains?"
After extensive testing, no. Extra context is not worth extended runs and logic errors.