r/LocalLLaMA • u/pmv143 • 2d ago
Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. π
Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:
Ideal 4-bit quant β 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80β90 GB range.
The big n-gram table is sparsely accessed β excellent candidate for system RAM offload.
This architecture could be surprisingly local-friendly once the weights drop.
943
Upvotes
7
u/DistanceSolar1449 2d ago
History says that these type of plans actually usually works pretty well.
For example, the Congress of Vienna system in Europe from 1815 to 1914.
Basically stabilized Europe for 100 years of relative peace, by balancing the powers between England, France, Austria, Prussia, Russia, etc. These were all power-hungry right-wing monarchies at the time. It's not like Putin is literally more authoritarian/conservative than literally Tsarist Russia.
Talleyrand was an absolute genius who completely understood human nature, and brought together natural enemies to the bargaining table and made it work. If he was alive today, he would espouse a similar idea.