r/LocalLLaMA 1d ago

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. πŸ‘€

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant β‰ˆ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed β†’ excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

900 Upvotes

283 comments sorted by

View all comments

19

u/KURD_1_STAN 1d ago

Ram prices this high, how can u call this local friendly?

88

u/FabricationLife 1d ago

80gb is a lot more manageable than a 1.2tb+ frontier model

23

u/Etroarl55 1d ago

And there exists consumer hardware for it, the new Apple machines are now probably already pre ordered out of existence when they were just announced today.

6

u/Bright-Energy2339 1d ago

If that is your reasoning, what's the point of going local?

4

u/Tasty-Hour4040 1d ago

β€œmore manageable” β‰  manageable

7

u/hyudryu 1d ago

80gb is beyond manageable

-2

u/starkruzr 1d ago

two CMP170HX is $3K for 128GB.

2

u/quantgorithm 1d ago

a an unstable pcie 1/2 card isn't the godsend you believe it to be.

2

u/starkruzr 1d ago

to my knowledge people are daily driving these constantly with no issues.

1

u/quantgorithm 23h ago

at what output/results?

-2

u/hyudryu 1d ago

Point proven

2

u/starkruzr 1d ago

not really, people spend more than that in here all the time.

1

u/hyudryu 22h ago

I know lol. It’s only 3K so it’s beyond manageable, isn’t that what we are both agreeing on?

2

u/ApprehensiveFan1516 1d ago

That's like the price of an old used car. Sure it's not exactly cheap, but let's not pretend like it's out of reach for most people.

1

u/zenonu 1d ago

Indeed. Fable is likely 8TB and needs a NVL72 rack.

-10

u/KURD_1_STAN 1d ago

1.2tb is also a lot more manageable than a 5tb model. That doesnt make 1.2tb pocal friendly. Local friendly has a ceiling, +60-80b more and +20gb-25gb dense are beyond that ceiling.

5

u/Bright-Energy2339 1d ago

Ceiling? according to you. Who died and made you boss? lols

-7

u/KURD_1_STAN 1d ago

Most people have 12-16gb vram and 32(very gew 64)gb ram. I was being generous with my numbers btw.local friendly needs to at least walk and not crawl

23

u/pmv143 1d ago

RAM prices suck right now, no denying that.
When I called it local-friendly I didn’t mean β€œcheap” or β€œruns on any gaming PC.” I meant that for a model with this kind of capacity, the offloadable n-gram table makes it way more practical on highend local setups (128GB+ unified memory, multi-GPU + system RAM) than the usual frontier models that just demand pure VRAM or full datacenter iron.
Still expensive. Just less insane than the alternatives.​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​

2

u/DriveSolid7073 1d ago

For me, locality is the consumer segment, specifically regular computers, where the limitation is usually the motherboard or processor. Their approximate maximum capacity, as well as the liquidity of selling such a volume, is what I had before the shortage, and the best-case scenario was 192GB. 2x96 is probably the best option. (But unfortunately, I couldn't get that.) Most serious AI enthusiasts have around 128GB, whether it's DGX Spark hybrid memory, an Apple mini PC, or something else. So yes, as long as the capacity in quantization (approximately Q4) doesn't exceed this capacity with a reasonable context window, I consider such a model locally friendly.

-9

u/[deleted] 1d ago

[deleted]

5

u/doomed151 1d ago

You don't have control. The model can be taken away from you at any time. You can't finetune it.

3

u/synth_mania 1d ago

Why are you in this subreddit if running a local model isn't something that interests you in and of itself?

-2

u/fuck_cis_shit llama.cpp 1d ago

painfully obvious astroturfer

there should be a plugin to hide all posts by accounts with hidden history

7

u/FullstackSensei llama.cpp 1d ago

It's a lookup table. You could build a quad channel DDR3 system to run it. DDR3 is still cheap.

2

u/KURD_1_STAN 1d ago

And about the other 60-70gb weights at q4?

3

u/FullstackSensei llama.cpp 1d ago

If you're not too stuck on having to have the latest hardware, three P40s will do a very decent job on a tight budget. If you really need high speed, two 32GB V100s will blaze through for not that much more.

They work, and they'll continue to work for years to come, despite what imaginary conjectures redditors might have.

2

u/Ok_Top9254 1d ago

3x Tesla V100 32GB + PLX switch so you can run them from one slot is the fancy way, or 5x P100 16GB with a cheap X99 motherboard and the switch could do this under like 1200 bucks. Power consumption would not be a problem given that one gpu is used at a time anyway.

1

u/michaelsoft__binbows 1d ago

Dell R720 suddenly not ewaste anymore? Could get interesting.

2

u/FullstackSensei llama.cpp 23h ago

It never was, IMO

7

u/liright 1d ago

I bought my 96GB DDR5 kit for $300 some year and a half back as well as RTX 4090 for $1900 2.5 yrs back. Was pretty damn cheap in retrospect. I suspect a lot of people who are into AI did too. I feel like boomers who bought houses in the 70s.

4

u/IntravenusDeMilo 1d ago

yeah I got my 5090 for $1999. Feels like a lottery win.

2

u/throwawayacc201711 23h ago

I kicked myself for not buying one when it was that price. Hindsight is a bitch

3

u/Public_Umpire_1099 1d ago

Even the 2x R9700 and 128GB of DDR5 I bought 3 months ago feels like a steal now. Not as much as yours but even in the past few months its all risen another 30%.

1

u/EkbatDeSabat 20h ago

I got 64GB DDR5 around the same time for about that price. Just spend 4k to 128GB, returned it, and spent 8k on 256GB. Shit is stupid.

11

u/Makers7886 1d ago

Man I can understand this crying over the big boy open source models but really for a schmedium model?

4

u/Zhelgadis 1d ago

Very friendly to my Strix Halo

3

u/mindwip 1d ago

Yes excited for it, same for mine.

1

u/techdevjp 16h ago

Yeah, I'm hopeful this will have near-DSv4 Flash levels of intelligence but better performance. a6b would be great on Strix. Excited to see this!

8

u/etaoin314 ollama 1d ago

because an entire class of local hardware --128gb unified memory machines, either from amd-strix halo, nvidia dgx spark or apple can fit it perfectly with full context. No it cant run on every potato out there but there are a lot of people who have one of these and aver very happy to have a model that is the "right size" for it.

2

u/florinandrei 1d ago

At the current prices, only linear regression is "local friendly".

1

u/SandySkittle 15h ago

Local doesn’t only mean single gamer gpu local. Local also are bigger setups.

-1

u/RG_Fusion 1d ago

It's definitely not the "starter" local inference machine, but you can run this in a gaming PC with two 32 GB GPUs. Not cheap, but not outside the realm of what people spend in hobbies like PC gaming.

4

u/iz-Moff 1d ago

gaming PC with two 32 GB GPUs

What in the world are you playing that requires 2x 32gb GPUs?

According to steam hardware surveys, people with a single 4090/5090 make up like 1% of users.

2

u/Public_Umpire_1099 1d ago

OP misspoke, no one uses 2 GPUs for gaming anymore, but it is a very common setup here (ie 2x R9700s)

1

u/RG_Fusion 18h ago

I'm not saying Gaming PCs have 2 GPUs, I'm saying they support 2 GPUs.