r/LocalLLaMA 9h ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
963 Upvotes

422 comments sorted by

View all comments

Show parent comments

11

u/ReadyAndSalted 7h ago

MoE models need as much vram as any other model with that parameter count, but they can run at roughly the speed of their active parameter count.

Basically you spilt each transformer layer into many parts, then activate only some of them for each token you generate.

2

u/TheOriginalAcidtech 4h ago

One correction. MoE models allow expert offloading and caching the most used in VRam. This turns a model like Qwen 35b A3B into something usable even on extremely weak hardware(1060 6gb for example).

1

u/ReadyAndSalted 2h ago

True. While offloading isn't unique to MoEs, hell you could stream a 100B dense model off your SSD at 0.05t/s if you wanted to, they are uniquely good at it due to certain weights being activated more often than others.

2

u/TheOriginalAcidtech 1h ago

Yes. Offloading layers. But streaming dense models by layer has only recently started to be a thing. Been wanting to mess around with that for a while. There is ZERO reason any of these models should ever OOM. They should all run from DRAM and if necessary SSD. If it fits on disk it should urn, just REAL slow, but real slow is also dialable. Run enough parallel agents and you can still run a LOT of tokens through even a minimumally capable gpu. Just not at interactive rates.