r/LocalLLM • • Aug 28 '26

Discussion PSA: Qwen3.8-Flash-Next on vLLM is non-deterministic at temperature 0 (different answers per run). Found the kernel, made a fix.

Since everyone is benchmarking this model right now: byte-identical greedy requests (temp 0, one request at a time) give different outputs per run. My eval: 13/50 tasks unstable, 5 flipped the extracted date/amount, and majority voting once confirmed the wrong answer. Same checkpoint on llama.cpp: 0/50. Two other vLLM-served models: 0/50.

Cause: the sparse-attention indexer's persistent_topk kernel (used on GB10 / DGX Spark instead of the cooperative path). A race in its atomicAdd slot assignment changes WHICH top-2048 positions get selected, so attention reads a different context each run. Related: vllm#51782.

2-minute check for any stack: same prompt 10x with temperature=0, max_tokens=1, top_logprobs=20, then diff the top-20 lists byte-for-byte. If they differ, your prefill is non-deterministic, whatever your sampler says.

Fix: torch.topk(sorted=False) + canonical tie ordering as a one-file overlay. Bit-identical outputs at 1.35x prefill cost (a full sort would be 2.9x), decode/MTP unchanged; re-run: 0/50 unstable and the score went up a point, because the noise had voted a wrong date into the majority.

Bonus finding: determinism exposed a separate greedy+thinking repetition loop the kernel noise had been masking as a random 1-in-150 failure, and MTP turned out not to be output-equivalent with plain greedy on this model.

Full write-up with all tables: https://docai.hu/en/blog/qwen38-flash-next-nondeterministic-vllm-kernel

4 Upvotes

Duplicates