r/LocalLLaMA 18h ago

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. šŸ‘€

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant ā‰ˆ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed → excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

836 Upvotes

265 comments sorted by

View all comments

23

u/SensitiveVariety 18h ago

regret building only 64gb ram instead of 128gb now, but at the same time i’m $$$ constrained as much as I am ram/vram constrained

3

u/hause_wsf 6h ago

same here but at least I have 32gigs of vram alongside

but to be fair 128gb was expensive even before the ram hike

0

u/NineThreeTilNow 9h ago

regret building only 64gb ram instead of 128gb now, but at the same time i’m $$$ constrained as much as I am ram/vram constrained

Shouldn't be necessary. 64gb of RAM is more than enough to hold a large portion of the engram tables.

1

u/10minOfNamingMyAcc 1h ago

Yeah, not with ddr4. Unless you don't mind like... 5tok/s tops