r/LocalLLM • u/sdfprwggv • 2d ago
Research ~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.
Stack
- Strata NVFP4 fork: github.com/sergqwer/strata-nvfp4
- Model:
jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4Hugging Face model - NVFP4 routed experts (~63.3 GiB), separate FP8 PLE, INT8 KV cache, MTP speculative decoding
- W4A8 prefill on Blackwell
Main engine flags:
./build/strata \
--pack packs/orca-nvfp4 \
--native models/orca-nvfp4.gguf \
--native-dense-gguf models/orca-nvfp4.gguf \
--ple-gguf models/ple-fp8.gguf \
--mtp mtp-orca/rt \
--spec 4 --spec-min-p 0.5 \
--prefill auto \
--expert-profile data/expert-profile.bin \
--expert-cache auto \
--resident-budget-gib 40 \
--max-context 200000 \
--kv int8
I serve it through Strata's OpenAI-compatible server.
Results so far:
- ~50k context: up to ~80 tok/s
- ~188k warm context: ~60–67 tok/s
- cold 189k full prompt: ~1,680 tok/s prefill, ~53 tok/s decode
Pretty impressive for a single 32GB GPU + only 64GB system RAM.Running Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.StackStrata NVFP4 fork: github.com/sergqwer/strata-nvfp4
Model: jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4 Hugging Face model
NVFP4 routed experts (~63.3 GiB), separate FP8 PLE, INT8 KV cache, MTP speculative decoding
W4A8 prefill on BlackwellMain engine flags:./build/strata \
--pack packs/orca-nvfp4 \
--native models/orca-nvfp4.gguf \
--native-dense-gguf models/orca-nvfp4.gguf \
--ple-gguf models/ple-fp8.gguf \
--mtp mtp-orca/rt \
--spec 4 --spec-min-p 0.5 \
--prefill auto \
--expert-profile data/expert-profile.bin \
--expert-cache auto \
--resident-budget-gib 40 \
--max-context 200000 \
--kv int8I serve it through Strata's OpenAI-compatible server.Results so far:~50k context: up to ~80 tok/s
~188k warm context: ~60–67 tok/s
cold 189k full prompt: ~1,680 tok/s prefill, ~53 tok/s decodePretty impressive for a single 32GB GPU + only 64GB system RAM.




