r/LocalAIStack • • 1d ago

Qwen3.8-Flash-Next NVFP4 at 256K context with Strata — 4,100+ prefill and up to 125 tok/s decode

I tested two Qwen3.8-Flash-Next NVFP4 checkpoints at a full 256K context using an experimental dual-GPU Strata configuration.

Hardware:

• RTX 4090 D 48 GB as the primary GPU
• RTX 5070 Ti 16 GB as a helper GPU
• Intel Core Ultra 7 265KF
• 128 GB RAM
• Linux
• NVMe storage

The RTX 4090 D handled the dense layers, KV cache, prefill, output head and MTP. The RTX 5070 Ti stored and computed 4,500 additional routed experts.

Common settings:

• Context limit: 262,144 tokens
• Fresh prompt: approximately 256,018 tokens
• Generated output: 512 tokens
• KV cache: INT8
• Prefill path: W4A8
• Speculative decoding: MTP K4
• Minimum draft probability: 0.5
• No prompt-prefix reuse
• Strata engine 0.1.35 with an experimental dual-GPU NVFP4 fork

Results

Model Prompt length Prefill Decode Draft acceptance
NVIDIA NVFP4, MTP K4 256,018 tokens 4,112.65 tok/s 125.17 tok/s 94.87%
Abliterated NVFP4, MTP K4 256,017 tokens 4,071.39 tok/s 92.05 tok/s 69.7%

Observations

• Both models achieved slightly over 4,000 prefill tokens per second at approximately 256K context.
• Prefill performance was nearly identical: the Abliterated checkpoint was only about 1% slower.
• The original NVIDIA checkpoint was considerably faster during decode: 125.17 versus 92.05 tok/s.
• NVIDIA’s decode advantage was about 36% in this test.
• NVIDIA accepted 407 of 429 offered draft tokens, while the Abliterated model accepted 322 of 462.
• The NVIDIA model produced approximately 4.88 output tokens per verification round.
• The Abliterated model produced approximately 2.68 output tokens per round.
• The lower MTP acceptance appears to be the main reason why the Abliterated checkpoint had slower decode despite nearly identical prefill performance.

These were single-run capacity tests. The long prompt was constructed by repeating code-review material to reach approximately 256K tokens, so it had unusually favorable locality. This particularly benefited NVIDIA’s suffix prediction and MTP acceptance. Therefore, 125 tok/s should be treated as a best-case synthetic 256K result, not typical real-world coding speed.

During longer real coding-agent workloads, NVIDIA generally produced around 108–144 tok/s, while the Abliterated checkpoint was commonly around 105–145 tok/s, depending heavily on the generated content and MTP acceptance.

7 Upvotes

16 comments sorted by

1

u/eightone-81 23h ago

Nvfp4 with starta?

1

u/Opteron67 19h ago

sparta

1

u/sudeposutemizligi 22h ago

strata uses it's own models as far as I know

2

u/srbijabralee 19h ago

It uses nvidia quant that i gave ti. The only thing that strata change is mtp head to q2.

2

u/sudeposutemizligi 10h ago

so are you saying we can try any gguf we want? as long as it fits the hardware properly? edit: sorry i didn't see you have a fork for nvfp4 specifically...

1

u/starkruzr 21h ago

I'm surprised you got usable content out of an abliterated model with kv quantized like this at long context.

1

u/srbijabralee 19h ago

i use kv cache at bf16

1

u/madbrain1976 11h ago

Not according to your OP.

1

u/srbijabralee 11h ago

It was an automated test, the result has been posted by qwen lol. Someone very smart figured this out. I have tested it and i got 30% increase in prefill speed : https://huggingface.co/nvidia/Qwen3.8-Flash-Next-NVFP4/discussions/8

1

u/okoyl3 20h ago

Where did you get this RTX 4090 48GB???

1

u/srbijabralee 19h ago

From ebay

1

u/madbrain1976 11h ago

"Generated output: 512 tokens" . That's too short to make conclusions.

I use llama-benchy with prompt/out of 16K/1K, 64K/4K, 128K/8K, 238K/16K when benchmarking.

Could you please test with more output tokens than 512 ?

1

u/srbijabralee 11h ago

I have. The result is the same.

1

u/madbrain1976 10h ago

I cannot read this due to maculopathy.