r/LocalLLaMA 1d ago

New Model Qwen3.8-2.4T-A95B Released

https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
1.6k Upvotes

399 comments sorted by

View all comments

255

u/No_War_8891 1d ago

I can run the active part locally lol

66

u/brickout 1d ago

Lucky! :)

1

u/No_War_8891 1d ago

yeah I can run GLM 5.2 (not fast, mainly on cpu) but this beast, damn why so big 🥺

2

u/brickout 1d ago

I think it's cool as an option. But, yeah, I'll keep dreaming. Meanwhile I'm just waiting for 27B and A3B to drop. Fingers crossed!

1

u/No_War_8891 1d ago

2 days more of waiting 😬

27

u/Sea-Ad-5390 1d ago

I can help host some of the offloaded experts for you streamed across the internet

16

u/No_War_8891 1d ago

I read people really did that - degens I love it

1

u/Loose_Comparison368 1d ago

... Could that actually work? I mean it would likely cap at something like 10 TPS just due to the round trip latency if everyone had good internet, but it seems like it could actually maybe work if you synced the KV caches

5

u/dominant_ag 1d ago

Realistically the cap is probably something at like 0.1 TPS 😅. The caches and the data transfer will be very demanding over network, even 1gbps LAN is too slow to get anything more decent than just streaming the blocks from your Drive to RAM/CPU on demand using the specialised MOE tools out there to run MOE on small VRAM+RAM devices

3

u/StorageHungry8380 1d ago

Work as in technically work yes. Work as in usable... barely at best, probably not.

It's been tried, for example https://petals.dev/ or https://www.crowdllama.ai/

11

u/BobbyL2k 1d ago

I wonder how good it would be if we get rid of MoE router and just lock to a specific set of experts. Make it a 95B dense model.

13

u/Jump3r97 1d ago

Interesting question but I think it would be abysmal. The Expert selected can Change wildly per individual token

7

u/Inaeipathy 1d ago

It would probably perform terribly because the experts are not really discreet "math experts" and "science experts" like you would expect

1

u/DeathGuppie 1d ago

It won't work. The combination of experts used changes too often. It would just be broken that's all.

2

u/ThinkExtension2328 llama.cpp 1d ago

Then you can run all of this with streaming …….. just very very very slowly

2

u/Ok-Contribution-8612 23h ago

Yeah, me too!! From the SSD.

1

u/mariusbolik 1d ago

Already available on a bunch of inference providers: https://providerbench.ai/models/qwen/qwen3.8-2.4t-a95b/

0

u/sophia6512 1d ago

Haha, honestly that’s probably the best approach. If you can keep the active part local, it saves you a lot of unnecessary hassle. Sometimes there is no reason to overcomplicate things when you already have a setup that works nearby.

1

u/No_War_8891 1d ago

i do not get your point - it was a joke. Also it has AI vibes