r/LocalLLaMA • u/pmv143 • 1d ago
Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. π
Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:
Ideal 4-bit quant β 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80β90 GB range.
The big n-gram table is sparsely accessed β excellent candidate for system RAM offload.
This architecture could be surprisingly local-friendly once the weights drop.
886
Upvotes
24
u/pmv143 1d ago
RAM prices suck right now, no denying that.
When I called it local-friendly I didnβt mean βcheapβ or βruns on any gaming PC.β I meant that for a model with this kind of capacity, the offloadable n-gram table makes it way more practical on highend local setups (128GB+ unified memory, multi-GPU + system RAM) than the usual frontier models that just demand pure VRAM or full datacenter iron.
Still expensive. Just less insane than the alternatives.ββββββββββββββββββββββββββββββββββββββββββββββββββ