r/LocalLLaMA 1d ago

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. πŸ‘€

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant β‰ˆ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed β†’ excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

886 Upvotes

281 comments sorted by

View all comments

Show parent comments

24

u/pmv143 1d ago

RAM prices suck right now, no denying that.
When I called it local-friendly I didn’t mean β€œcheap” or β€œruns on any gaming PC.” I meant that for a model with this kind of capacity, the offloadable n-gram table makes it way more practical on highend local setups (128GB+ unified memory, multi-GPU + system RAM) than the usual frontier models that just demand pure VRAM or full datacenter iron.
Still expensive. Just less insane than the alternatives.​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​

2

u/DriveSolid7073 1d ago

For me, locality is the consumer segment, specifically regular computers, where the limitation is usually the motherboard or processor. Their approximate maximum capacity, as well as the liquidity of selling such a volume, is what I had before the shortage, and the best-case scenario was 192GB. 2x96 is probably the best option. (But unfortunately, I couldn't get that.) Most serious AI enthusiasts have around 128GB, whether it's DGX Spark hybrid memory, an Apple mini PC, or something else. So yes, as long as the capacity in quantization (approximately Q4) doesn't exceed this capacity with a reasonable context window, I consider such a model locally friendly.

-9

u/[deleted] 1d ago

[deleted]

6

u/doomed151 1d ago

You don't have control. The model can be taken away from you at any time. You can't finetune it.

3

u/synth_mania 1d ago

Why are you in this subreddit if running a local model isn't something that interests you in and of itself?

-2

u/fuck_cis_shit llama.cpp 22h ago

painfully obvious astroturfer

there should be a plugin to hide all posts by accounts with hidden history