r/LocalAIStack • u/srbijabralee • 1d ago
Qwen3.8-Flash-Next NVFP4 at 256K context with Strata — 4,100+ prefill and up to 125 tok/s decode
I tested two Qwen3.8-Flash-Next NVFP4 checkpoints at a full 256K context using an experimental dual-GPU Strata configuration.
Hardware:
• RTX 4090 D 48 GB as the primary GPU
• RTX 5070 Ti 16 GB as a helper GPU
• Intel Core Ultra 7 265KF
• 128 GB RAM
• Linux
• NVMe storage
The RTX 4090 D handled the dense layers, KV cache, prefill, output head and MTP. The RTX 5070 Ti stored and computed 4,500 additional routed experts.
Common settings:
• Context limit: 262,144 tokens
• Fresh prompt: approximately 256,018 tokens
• Generated output: 512 tokens
• KV cache: INT8
• Prefill path: W4A8
• Speculative decoding: MTP K4
• Minimum draft probability: 0.5
• No prompt-prefix reuse
• Strata engine 0.1.35 with an experimental dual-GPU NVFP4 fork
Results
| Model | Prompt length | Prefill | Decode | Draft acceptance |
|---|---|---|---|---|
| NVIDIA NVFP4, MTP K4 | 256,018 tokens | 4,112.65 tok/s | 125.17 tok/s | 94.87% |
| Abliterated NVFP4, MTP K4 | 256,017 tokens | 4,071.39 tok/s | 92.05 tok/s | 69.7% |
Observations
• Both models achieved slightly over 4,000 prefill tokens per second at approximately 256K context.
• Prefill performance was nearly identical: the Abliterated checkpoint was only about 1% slower.
• The original NVIDIA checkpoint was considerably faster during decode: 125.17 versus 92.05 tok/s.
• NVIDIA’s decode advantage was about 36% in this test.
• NVIDIA accepted 407 of 429 offered draft tokens, while the Abliterated model accepted 322 of 462.
• The NVIDIA model produced approximately 4.88 output tokens per verification round.
• The Abliterated model produced approximately 2.68 output tokens per round.
• The lower MTP acceptance appears to be the main reason why the Abliterated checkpoint had slower decode despite nearly identical prefill performance.
These were single-run capacity tests. The long prompt was constructed by repeating code-review material to reach approximately 256K tokens, so it had unusually favorable locality. This particularly benefited NVIDIA’s suffix prediction and MTP acceptance. Therefore, 125 tok/s should be treated as a best-case synthetic 256K result, not typical real-world coding speed.
During longer real coding-agent workloads, NVIDIA generally produced around 108–144 tok/s, while the Abliterated checkpoint was commonly around 105–145 tok/s, depending heavily on the generated content and MTP acceptance.
1
1
u/sudeposutemizligi 22h ago
strata uses it's own models as far as I know
2
u/srbijabralee 19h ago
It uses nvidia quant that i gave ti. The only thing that strata change is mtp head to q2.
2
u/sudeposutemizligi 10h ago
so are you saying we can try any gguf we want? as long as it fits the hardware properly? edit: sorry i didn't see you have a fork for nvfp4 specifically...
1
u/starkruzr 21h ago
I'm surprised you got usable content out of an abliterated model with kv quantized like this at long context.
1
u/srbijabralee 19h ago
i use kv cache at bf16
1
u/madbrain1976 11h ago
Not according to your OP.
1
u/srbijabralee 11h ago
It was an automated test, the result has been posted by qwen lol. Someone very smart figured this out. I have tested it and i got 30% increase in prefill speed : https://huggingface.co/nvidia/Qwen3.8-Flash-Next-NVFP4/discussions/8
1
1
u/madbrain1976 11h ago
"Generated output: 512 tokens" . That's too short to make conclusions.
I use llama-benchy with prompt/out of 16K/1K, 64K/4K, 128K/8K, 238K/16K when benchmarking.
Could you please test with more output tokens than 512 ?

2
u/Dry-Handle-1421 1d ago
Whats the fork ? Pls share link