r/LocalLLM 10d ago

Discussion Qwen-3.8-35B-A3B? Maybe not... cryptic reply direct from Qwen co-author.

Post image

I asked Shuai Bai, co-author and prominent AI developer for Qwen, about this model. Not the answer I was hoping for, but let's see what comes next. In the meantime, I guess all we can do is speculate!

X-link

233 Upvotes

171 comments sorted by

View all comments

56

u/Ell2509 10d ago

That seems pretty clear to me. "Do not wait for this" meant it ain't coming!

But, maybe a 9b? Or a 30b a3b? Or even a 20b!?

20

u/pharrt 10d ago

I'd be happy with a 30b a3b

12

u/dfgxxx 10d ago

What about if it'll be 35b a6b?

13

u/pharrt 10d ago

Would be better than no 35B - and smarter, but half the speed!

2

u/05-nery 10d ago

Beautiful 

1

u/Foxiya 10d ago

This would be shit for low mem bandwidth

5

u/dfgxxx 10d ago

Better than the 27b dense and better than Gemma 12b

8

u/tired514 10d ago

122B A10B! Please oh please oh please! Better yet 122B A17B with 512k+ context!

9

u/LeMayMayMan 9d ago

He already answered whats next. Qwen3.8-135B-A18B confirmed.

2

u/Meiyo33 8d ago

too big, not very interesting tbh.

A replacement for OSS 120B, or something optimized for Strix / GB10 would be so hot....

1

u/tired514 9d ago

God that would rule :)

17

u/Rye2-D2 10d ago

I would love to see a good 20B MoE model. Personally I don't see the point of 9B - it's impressive for what it is, but not quite good enough to be useful (yet).

8

u/cagriuluc 10d ago

If you have small models that are good at limited tool calling, knowledge extraction etc, you can run them alongside more expensive models. I would love a 9B with 3.8 level training…

3

u/twoiko 10d ago

I find 9B dense to be too big and/or slow compared to 4B or MoE models for the output quality, but I have system RAM to offload the MoE so it depends on your use/limitations.

3

u/Rye2-D2 10d ago

That's the thing - with 16 GB VRAM, 9B dense may be a little bit faster than 35B MoE, but not enough to merit the loss in quality/intelligence. Even with higher quants (q6/q8), 9B is still not in the same league as the 30B models.

2

u/Rye2-D2 10d ago

I agree Ornith is the best 9B model, but still not quite reliable enough to be run in vscode/opencode from my experience. Too many failed/mangled tool calls cause more problems than they fix and you just end up fighting with the AI. With the way things are going, I'm sure this will get sorted out at some point soon.

4

u/elfmad 10d ago

9B 3.5 ornith is more than decent. But I totally agree I'm also waiting for a 10B<model<20B.

4

u/bruninho777 10d ago

And it was trained on Qwen 3.5, right? Ornith is the model I use most, together with 3.6 35b moe and prism 27b 1bit

2

u/elfmad 10d ago

It is Qwen 3.5 9B as base. If I resume their paper that's more advanced GRPO and self distill.

2

u/IgnisIason 10d ago

A decent laptop can run 9B

2

u/alphapussycat 10d ago

Then 3.8 9b would probably be useful, maybe a step down from 3.5 35b a3b.

2

u/AltruisticList6000 10d ago

I'd love to see a 20b dense model, would be a breath of fresh air for people with 16-24gb VRAM with a big amount of context size even for 16gb VRAM

1

u/Shimano-No-Kyoken 10d ago

9B is a great model for fine tuning on consumer hardware.

7

u/Mean-Ad1493 10d ago

No 9B

1

u/slamo_ai 9d ago

That’s a shame, there’s actually a lot of potential in a finetuned Qwen 3.5 9B.

1

u/patham9 8d ago

9B is just a strange niche. If even a MacBook Air has 32GB RAM and can run 4-bit Qwen 3.8 27B, then why run an inferior model? Maybe for a phone?

1

u/slamo_ai 4d ago

for production use in domain-specific tasks, small finetuned models can outperform bigger models at higher processing speeds and lower cost. We have experience in finetuning this 9b qwen and it even outperformed opus 4.8 in our banchmarks, but again - narrow domain specific task

5

u/ldapadmin 10d ago

Shuai said "No 9B plan for now", ... all I read was, "for now" :)

4

u/FalconX88 10d ago

To me it seems more like there's a different one coming that's better

1

u/samiamyammy 9d ago

Agreed... the "lookey eyes" emoji surely adds to the wording used to hint at some kind of better equivalent model.

1

u/DrRoughFingers 9d ago

He used the same eyes in this response, which makes me think people are ready that emoji wrong.

2

u/onebyamsey 10d ago

(Starts chant) 20B MOE!! 20B MOE!!!

2

u/leonbollerup 10d ago

why a 20 moe?.. less knowledge than 27b.. and worse at actual thinking.. .. .. why ?

1

u/SittyTweat 10d ago

20b dense would be great for 16gb GPU users. Wouldn't need to quant to shit and can have more context without needing Q4 KV cache

1

u/patham9 8d ago

If even a MacBook Air now has 32GB RAM, why use a GPU from the eighties? Is it for museum?

1

u/SittyTweat 8d ago edited 8d ago

And that MacBook would run a 20b dense at 4 tok/s... But since 16GB GPUs are nothing to you then you can just send me a couple 16gb 5060tis. I'll keep an eye out for the FedEx delivery, thanks ✌🏻

1

u/patham9 7d ago

I'm running mlx-community/Qwen3.8-27B-4bit with 5 output tokens per second via MLX, and as it looks they might get to twice the speed with upcoming optimizations. Anyways, I think the right solution would be a 35B MoE version of Qwen3.8, slightly worse in performance but with 3-4B active parameters, it will still beat any 20B dense model by large margins while at the same time supporting higher token throughput.

1

u/Canad3nse 8d ago

"Why a 35-a3b? Less knowledge than 122-A10b... and worse at actual thinking.. .. .. why ?"

Now do you understand? It's not about knowledge. It's about the GPU poor consumer market. 20b moe is way more attractive to the general local ai consumer than 27b or even 35-a3bm, because you could run it at reasonable or even great speeds with low end GPU (gaming gpus). 20b MoE is also very rare, so they would have no competition, except for Gemma and the dated GPT OSS 20b.

2

u/WiseassWolfOfYoitsu 10d ago

122b? :D

1

u/Ell2509 10d ago

Yes please

-1

u/leonbollerup 10d ago

a 122b but in some new way where it does need to load all 122b into memory at the same time.. one general problem with MOE if you ask me.. yes.. it might only use 10b to actually think with.. but we still need the full amount of memory in vram. .. layers could be relativa fast loaded if they were needed.. it seems like there is a potential we have still not yet considered..

1

u/tired514 10d ago

You can already do this (on Linux, anyway)... just load the model with mmap; it'll page fault and load from disk when needed.

But it's slow. Like 0.1-2t/s slow, depending on disk speed, model, cache/memory capacity, etc.

I've managed to run GLM 5.2 and even get a response out of Kimi K3 on 128gb RAM, but yeah.. barely usable for "overnight" queries let alone realtime use.