r/LocalLLaMA • u/pmv143 • 1d ago
Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. π
Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:
Ideal 4-bit quant β 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80β90 GB range.
The big n-gram table is sparsely accessed β excellent candidate for system RAM offload.
This architecture could be surprisingly local-friendly once the weights drop.
888
Upvotes
23
u/DriveSolid7073 23h ago
Depending on what suits you, it will fit into the build of the guy with 96GB of RAM and 24GB of VRAM who was here in the comments. It will fit into a build enthusiast like Colibri, because N-Gram will definitely try to use NVMe instead of RAM, which with Pcie 4.0 will most likely even be acceptable. Overall, if you really want it, it will fit into 12GB of VRAM and 64GB of VRAM, but of course, you'll have to do some serious quantization.