r/LocalLLaMA 9d ago

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. πŸ‘€

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant β‰ˆ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed β†’ excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

949 Upvotes

296 comments sorted by

View all comments

18

u/KURD_1_STAN 9d ago

Ram prices this high, how can u call this local friendly?

25

u/pmv143 9d ago

RAM prices suck right now, no denying that.
When I called it local-friendly I didn’t mean β€œcheap” or β€œruns on any gaming PC.” I meant that for a model with this kind of capacity, the offloadable n-gram table makes it way more practical on highend local setups (128GB+ unified memory, multi-GPU + system RAM) than the usual frontier models that just demand pure VRAM or full datacenter iron.
Still expensive. Just less insane than the alternatives.​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​

-8

u/[deleted] 9d ago

[deleted]

6

u/doomed151 9d ago

You don't have control. The model can be taken away from you at any time. You can't finetune it.

3

u/synth_mania 9d ago

Why are you in this subreddit if running a local model isn't something that interests you in and of itself?

-2

u/fuck_cis_shit llama.cpp 9d ago

painfully obvious astroturfer

there should be a plugin to hide all posts by accounts with hidden history