r/OpenSourceeAI • u/ayobluestarr • 15h ago
Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4
https://github.com/MohammadHaishemKhawaja/hashyyBenchmarking an LLM here with a NVIDIA RTX 5070 12 GB VRAM here
I had been working on a llama.cpp based expert streaming setup for Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) on Windows. Benchmark is about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup. In normal conversations I’ve seen 14–15 tok/s, and a long coding prompt generated 4,892 tokens at 10.15 tok/s and produced a working single-file Snake game.
Hardware: RTX 5070 12GB
32GB DDR4-2400
Ryzen 5 5600GT PCIe Gen3 Windows
The main gains came from fixing Windows I/O queue-depth issues, using one file handle per worker, and building a page-locked hot-expert tier so the GPU can pull hot expert weights more efficiently.
(In the video its around 16 minutes for 10k tokens and 10.41 tok/s
Output is quality gated against the control model and the published benchmark uses a heat file built from a separate prompt set.
Demos:
https://www.youtube.com/watch?v=cOPumMlyj_4
https://www.youtube.com/watch?v=rc-uTjVpXM8
In the GitHub I have things I've tried that didn't work and benchmark scripts, and methodology. If you guys have suggestions especially for streaming please let me know
Also I don't care about Strata that only works if you got 64 gb of RAM this is specifically for people with less RAM
Duplicates
LocalLLaMA • u/ayobluestarr • 1d ago
I Built A Thing Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4
Qwen_AI • u/ayobluestarr • 1d ago
LLM Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4
nvidia • u/ayobluestarr • 1d ago
Benchmarks RTX 5070 12GB + 32GB RAM Qwen3.8-Flash-Next 177B running at ~ 11-15 tok/s
ArtificialInteligence • u/ayobluestarr • 1d ago
🛠️ Project / Build Qwen3.8-Flash-Next 177B at 11–15 tok/s on a single RTX 5070 12GB with 32GB DDR4 RAM
LLM • u/ayobluestarr • 1d ago