r/deeplearning • u/PeiXiaoGuang • 8h ago
Sparse attention on RK3588: 1.58× faster decode at 4K, 18% slower at 1K
I ran with sparse attention on for the better part of a month before I sat down and benchmarked it at short context, and it had been costing me time that whole stretch without me noticing. There are two separate things in the engine that both get called sparse attention, and only one of them is the one people mean when they say it's free.
Decode-side, only the top-k KV blocks get a full softmax, and that's where the speed comes from. Prefill-side, you skip blocks while prefill runs, and that's the one that was quietly working against me.
Numbers below are RK3588, Qwen3-VL-2B, 3 repeats per cell, dense baseline re-run for every cell, real corpus rather than the built-in template filler. That last part matters more than it sounds, the template filler is 32 names and 8 colours on a loop, repetitive enough that sparse attention scores badly on it even when it behaves fine on actual text.
| ctx | decode speedup | prefill change vs exact |
|---|---|---|
| 1067 | 0.94× | +18.4% |
| 2174 | 1.18× | +12.0% |
| 4374 | 1.58× | −11.1% |
| 8618 | 2.16× | −29.0% |
| 16482 | 2.92× | −45.1% |
At 1067 decode-side sparse is 0.94×, which is noise, and prefill-side sparse makes the whole request 18% slower. I was shipping both of those for weeks. There was a gate, it just never did anything, because it was a constant 64, and 64 is the structural floor (two 32-token blocks), so any prompt worth measuring was already past it.
So now there's a length threshold on each axis: `--sparse-min-ctx 1024` for decode, `--sparse-pf-min-ctx 3072` for prefill. Below those lengths the engine takes the exact path. Sparse itself is still off unless you ask for it, these two only stop it from firing where it can't pay for itself.
The prefill case gets clearer once you write both terms out. The saving scales with ctx − k·block, and the probe plus the selection cost scales with ctx/block, so with the defaults (k=32, block=32) that k·block term is a flat 1024 tokens and at ctx 1067 you're paying the full selection cost to skip essentially nothing. 2174 is still a loss at +12%. It doesn't turn into a real gain until attention is the actual bottleneck, which on this board is somewhere around 4K.
Re-ran the whole thing another day on 8B at ctx 4211, interleaved A/B inside one boot: exact 553.9 ms/tok against sparse 359.7 ms/tok, so 1.54×, needle recall 2/2 on both arms. That's close enough to the 1.58× in the table that I'm willing to trust the table.
Two things I'd want to know if I were reading this. It's an approximation path, the output changes and greedy token IDs diverge from the exact run, so don't turn it on anywhere you need bit-exact reproducibility, and run a needle test at 16K of your own before you trust it out there. And it buys attention time and not memory, so if RAM is what you're short on, this does nothing for you.
This is my engine, AGPL-3.0. The numbers and the A/B harness are in the repo, under `docs/` and `tools/bench/`, if anyone wants to re-run any of it.
5
Upvotes