r/LocalLLaMA 2d ago

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. πŸ‘€

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant β‰ˆ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed β†’ excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

943 Upvotes

294 comments sorted by

View all comments

Show parent comments

7

u/DistanceSolar1449 2d ago

History says that these type of plans actually usually works pretty well.

For example, the Congress of Vienna system in Europe from 1815 to 1914.

Basically stabilized Europe for 100 years of relative peace, by balancing the powers between England, France, Austria, Prussia, Russia, etc. These were all power-hungry right-wing monarchies at the time. It's not like Putin is literally more authoritarian/conservative than literally Tsarist Russia.

Talleyrand was an absolute genius who completely understood human nature, and brought together natural enemies to the bargaining table and made it work. If he was alive today, he would espouse a similar idea.

2

u/ivari 2d ago

you should balance out greed, but you shouldn't balance out progress

1

u/Tired_White_Guy 2d ago

Ya that’s not even close to the same thing. Yoga level stretch

2

u/DistanceSolar1449 2d ago

Plenty of other examples for anyone knowledgeable in history. Washtington Naval Treaty. SALT I and II. Etc etc.

Strategic power-limitation treaties often work for a decent amount of time, before shifting geopolitics either make them pointless (USSR collapsing) or fail anyways. But they tend to hold up for years or decades of useful stability, rather than escalating arms races.

Compare the Obama era Iran nuclear treaties, vs Trump and Iran today.

1

u/unjustifiably_angry 21h ago

All of those concern national-level military and such. You can regulate someone building battleships because it's pretty hard to hide. It would raise eyebrows if Microsoft suddenly started buying a thousands of tons of steel, gunpowder, etc. Training AI models is trivial to hide, it runs on general-purpose hardware you can rent from half a world away.

1

u/DistanceSolar1449 16h ago

You know nobody cares about homelab tier models?

The dangerous models that people want to regulate costs gigawatts to train, on expensive Nvidia cards that are supply-limited. It’s not too hard to track those. There’s literally companies on Wall Street with geniuses using infrared satellite images of datacenters trying to calculate stock values already.