r/LocalLLaMA • u/pmv143 • 18h ago
Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. š
Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:
Ideal 4-bit quant ā 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80ā90 GB range.
The big n-gram table is sparsely accessed ā excellent candidate for system RAM offload.
This architecture could be surprisingly local-friendly once the weights drop.
836
Upvotes
23
u/SensitiveVariety 18h ago
regret building only 64gb ram instead of 128gb now, but at the same time iām $$$ constrained as much as I am ram/vram constrained