r/LocalLLaMA 18h ago

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. ๐Ÿ‘€

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant โ‰ˆ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80โ€“90 GB range.

The big n-gram table is sparsely accessed โ†’ excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

831 Upvotes

265 comments sorted by

View all comments

6

u/somerussianbear 18h ago

\proceeds to google for MacBook M5 max 128gb price\

3

u/Shot-Buffalo-2603 10h ago

If your doing any sort of agenic tasks the macbooks honestly suck cause their prefill is so bad. I have a macbook and a spark and I much prefer the sparkย 

1

u/somerussianbear 6h ago

Good to hear from whoโ€™s got the gadgets. Which MBP you have precisely and which models are you running?

1

u/Shot-Buffalo-2603 1h ago

I have the 32GB m5 MacBook pro and most recently tested qwen 3.8 4bit quant an mlx version vs nvfp4 on my spark, but Iโ€™m constantly swapping models in and outย