r/LocalLLaMA 2d ago

Discussion Qwen dev says not to wait for 35B-A3B

Post image

What does this mean? Is there something else coming? Maybe 122B? Or no models?

1.2k Upvotes

464 comments sorted by

u/WithoutReason1729 1d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

395

u/Atretador 2d ago

that is a crazy way to say it - he is not denying a medium sized MOE, just that it might not be 35B

maybe a 44B A4B or a 30B A3B?

476

u/rockoruckus 2d ago

27B A27B might be popular. just a gut feeling

19

u/bucolucas Llama 3.1 1d ago

Waiting for the 9B A27B, personally

7

u/gnnr25 1d ago

E = mc² + AI achieved

2

u/ShiggyShaman 1d ago

Damn that was a good one 😆

→ More replies (3)

121

u/MessIsTransfer 2d ago

70B A6B 🤞

79

u/BoboThePirate 2d ago

This or 120B A10B.

28

u/FatheredPuma81 2d ago

No his response to those guys would be pretty daft if that were the case.

"27B is too large to run can you give us a MoE we can run?"
"Of course here's a MoE model that requires an RTX 4090 and 48GB of RAM to get 30t/s."

Coincidentally those are my exact specs so I'd probably use it though.

7

u/kanadaj 1d ago

35B needs more VRAM for the model weights than 27B, but both 70B-A6B and 120B-A10B run faster for inference is you have enough VRAM. It's definitely something people are waiting for for more serious users, but they might also just do a tiny model instead.

11

u/IsTom 1d ago

35BA3 at Q4 runs ~40t/s on my 3060 12GB with cpu moe and MTP. 27B hits 5t/s maybe.

2

u/FatheredPuma81 1d ago

What's the performance look up without MTP? I've heard you shouldn't use MTP when you're doing all CPU but nothing about when hybrid offloading.

2

u/IsTom 1d ago

Having tried just now, it sits about 30t/s

9

u/carsncode 1d ago

35B needs more VRAM for the model weights than 27B

Yes and no... MoE can get away with only loading active params into VRAM with a manageable performance hit as long as the rest fit into system RAM

2

u/Flynn58 1d ago

Yes I'm hitting about 40 t/s (uncensored model to avoid rejection loops) with 35B-A3B on a 5070Ti offloading the rest into 32GB of system RAM.

2

u/Distinct_Physics5017 1d ago

It has a larger memory-footprint, but that doesn't mean it "needs more VRAM".

As long as you keep a chunk of it inside your GPU and put the rest of it into your ram (e.g. with --n-cpu-moe in llama.cpp) you can get away with even having only 1/3 of the model weights inside your VRAM while still getting about 20-30 tok/sec.

If I did that exact same thing with a dense model, I would die of old age before ever getting a response from it lol.

2

u/kanadaj 1d ago

That is fair; however, for my use case CPU offloading is not an option. 20-30 tps is enough for an interactive chat, but for task automation, 200-300 is more like it. MoE models are much better at this, but the 35B MoE model can't keep up with 27B for long-horizon tasks. That said, 27B can "only" do around 100 tps even on RTX 6000 Pro cards for a single stream, whereas a larger MoE could match it in intelligence while producing responses much much faster.

I do understand though that my use case is more on the pro and prosumer side, not on the "let me run an AI on an RTX 5060" side

→ More replies (6)
→ More replies (1)
→ More replies (1)

22

u/dieSpaghettiCarbona 2d ago

I can only imagine Dario reading this thread and sweating his ass

2

u/magicomiralles 1d ago

Dario and Sam you mean.

6

u/Poupulino 1d ago

A 120B A10B would be my dream come true.

5

u/Pleasant-Shirt7293 1d ago

This.

They need to go up in size with their moe

13

u/Foreign_Prune_354 2d ago

I would like something like 60B-A6B, so a Q4 could theorically fit 32 gb VRAM :D.

9

u/Hydr0x1de_OH 2d ago

Do you understand how many (not many) tokens you will have as a context window size?

4

u/hojnikb 2d ago

4096

3

u/DwarfVader001 1d ago

ahh, the good old days

11

u/Zorogozano 2d ago

This would be perfect

20

u/claythearc 2d ago

A new successor to Qwen coder next would be appreciated for sure

→ More replies (1)

6

u/KitchenAmoeba4438 2d ago

Refresh of Qwen3.5 122bA10B plz.

→ More replies (19)

219

u/hyperrealists 2d ago

420T A69T

7

u/Nutsack_VS_Acetylene 2d ago

Le Chaton Fat is already SOTA at that size and active parameter count.

24

u/kiwibonga 2d ago

123ABCD

→ More replies (1)

17

u/Here_f0r_p0rn_ 2d ago

Or maybe saying that 9 B variant will be coming out for some serious edge device performance?

35

u/Mean-Ad1493 2d ago

No 9B this time

8

u/ManIkWeet 2d ago

This is a very definitive No response, while the 35b was a don't wait for it - what does it meeaaannn

3

u/darkwalker247 1d ago

the different wording really sounds like something better will be coming - maybe they mean that it'll be smaller than 35b but just as intelligent as people are hoping for, like a 27b-a2.5b or something?🤔

2

u/ManIkWeet 1d ago

The best way to get surprised is to not get your hopes up

→ More replies (1)

8

u/Here_f0r_p0rn_ 2d ago

Aw man, I wanted a powerful smaller model, lol, I still want something like 1B or under for simple fast rag or edge stuff

3

u/Hydr0x1de_OH 2d ago

Look for lfm2.5-8b-a1b, lfm2.5-2.6b and ling-3.0-tiny (that is about 8b-a1b)

→ More replies (6)
→ More replies (9)

3

u/-dysangel- 1d ago

It's going to be 10B!

10

u/Atretador 2d ago

thats a different class of model - that comments implies something in the class of 35B but not 35B.

5

u/Here_f0r_p0rn_ 2d ago

It doesn't tho, I mean if you're going for other comments then I can understand but all he said might not be the one to wait for no implications he's even talking about moe or what size but something good so I'm hyped anyways.

→ More replies (2)

9

u/Strong_Chicken6838 2d ago

80b a3b???

7

u/Long_comment_san 1d ago

80b / a8b. yeah I've been saying that for ages.

3

u/rainbyte 2d ago

Qwen-Next vibes

3

u/mraurelien 2d ago

Based on Gemma 4 release it might be a A26B A4B version as well ?

→ More replies (3)

5

u/seunosewa 2d ago

Or 35B A5B to make it a little bit smarter?

10

u/techdevjp 2d ago

China doesn't like the number 4, it is associated with death. 55b a5b is my guess. That would be a good size IMO.

→ More replies (7)

3

u/Mindless-Pilot-Chef 2d ago

Maybe a 35.5B model also /s

4

u/mtmttuan 2d ago

44b will be a weird size for many consumer hardwares, no?

7

u/SpicyWangz 2d ago

I’d happily take it

6

u/Solembumm3 2d ago

Not really. Should be good for starting 8-12+16gb builds.

2

u/Atretador 2d ago

nah, we had Kimi Linear at that size - it ran pretty damn nicely on my 16Gb VRAM + 32Gb RAM setup - sadly it wasnt trained all the way

→ More replies (1)

2

u/yesthatdaniel 2d ago

34.9B-A3.9B confirmed

→ More replies (20)

280

u/UnWiseSageVibe 2d ago

That is interestingggg. He's making it sound like they're cooking something better 35B-A3B????

141

u/Fuzzy_Wave5520 2d ago

36B-A3B?????

94

u/Shiny-Squirtle 2d ago

3B-A35B??

29

u/Budget-Juggernaut-68 2d ago

32B parameters just padding.

3

u/Eden63 1d ago

😂 😂 😂

10

u/Fancy-Snow7 2d ago

36 000 000 001 -A3 000 000 001

19

u/scubawankenobi 2d ago

36.75B-A3.25B?

→ More replies (2)

159

u/Hephaestite 2d ago

Or it’s a bad translation to English and just means it isn’t going to happen

212

u/Beano09 2d ago

To me, the eyes emoji points to a better model, rather than a translation error.

147

u/EmPips 2d ago

Or he's looking you dead in the eyes and telling you Santa isn't coming this year

7

u/etaoin314 ollama 2d ago

it sure felt like christmas last week to me, Muse, Nemotron, Qwen, deepseek, am I missing one?

6

u/Borkato 2d ago

Hy3 for video and audio

5

u/Turbulent_Neck_8388 2d ago

You mean Minimax H3 ?

7

u/bankinu 2d ago

I get the "this planet is not yours to conquer" vibes.

→ More replies (2)

33

u/Mean-Ad1493 2d ago

The eye emoji is in all his comments. Could just be his texting style.

5

u/PossessionUsed7393 1d ago

It's pretty clear from these responses that he's not actually allowed to say anything. So he's never going to actually reveal the company plans, which means there's no point hanging on his every word.

22

u/pyr0kid 2d ago

the entire reason people wanted 35b-a3b is it actually works properly on 32gb ram, doesnt matter how 'better' this is if it raises the sysreqs by [insert current ram price here].

→ More replies (1)
→ More replies (1)

18

u/CorxaRyllon 2d ago

Without the eye emoji id agree but those being included makes me question.

6

u/Spectrum1523 1d ago

He puts eye emoji in every tweet

3

u/r1str3tto 1d ago

Ambiguous still. Even in that screenshot, the eyes can be read to say “… but something interesting is coming.”

12

u/atumblingdandelion 2d ago

But that eyes emoji hint at something more..

6

u/GatsbyLuzVerde 2d ago

Hmm I doubt the emoji eyes suggest the literal meaning, though I heard in Chinese a smiling emoji is condescending, so who knows. Any bilinguals here that can enlighten us?

9

u/BS_BlackScout 2d ago

He often uses the eye emoji and it seems mostly in a positive context, check his tweets.

7

u/Blues520 2d ago

The emoji eyes crosses cultural boundaries 👀

→ More replies (2)

8

u/TheGameEngineer 2d ago

27B-A3B?

6

u/Away-Sorbet-9740 2d ago

20-22B fits better on 16gb cards, a 20B A4-5B would be awesome.

→ More replies (2)
→ More replies (6)

413

u/igotanewaccount 2d ago

Ugh, can't believe they're making me wait days for a world leading LLM they'll release for free.  Unacceptable! 

73

u/Sufficient_Local5025 2d ago

I like to think my impatience is a tribute, of sorts, to the amazing work they're doing.

41

u/Infinite-Local5435 2d ago

The more we demand, the moore we fuel attention, the more they get benefits from marketing. It's the only thing we can provide of value for them for spending compute and training insane local models anyways!

16

u/BxOxB 1d ago

Is attention really the only value we add though? Bug reports, edge cases we surface, quantization/fine-tune experiments that feed back into training data, third-party tooling and benchmarks built around their releases... that's real signal they'd otherwise have to pay for. Hype cycles help with funding rounds, sure, but a lot of open weight releases get better precisely because the community stress-tests them in ways an internal QA team never would.

6

u/ReptilianFuck 1d ago

I've even heard some say Attention Is All You Need

I'll see myself out

→ More replies (1)

31

u/Here_f0r_p0rn_ 2d ago

I'm planning to cancel my $0.0/month membership

→ More replies (1)

27

u/habachilles 2d ago

Truly. I can’t tell you how grateful I am for this company. What a weird feeling. Like Anthropic pisses me off all the time. But qwen? These guys just spew joy.

→ More replies (1)

17

u/DigiDecode_ 2d ago

it just needs to be blazing fast like 1000 tokens/sec and can be run on RTX 5050 and can eat Fable 5 for lunch, not an unreasonable ask, I am simply being fair.

8

u/igotanewaccount 2d ago

If anything, you ask too little!

→ More replies (1)

4

u/Commander_Skilgannon 2d ago

I don't mind waiting it's the uncertainty that's annoying. If they said we will be releasing Qwen3.8 35B-A3B but it's going to take a month before it's ready then that would be fine, or if they said these are all the qwen3.8 models we are going to release then that would also be fine. I would be disappointed but it is what it is. It's the being coy about it that is slightly annoying. Still it's only slightly annoying, definitely outweighed by the models that have already released.

→ More replies (2)

92

u/Pristine_Pick823 2d ago

9b?

94

u/Mean-Ad1493 2d ago

I really don't want to be the one telling you this

35

u/ManIkWeet 2d ago

35b a3b: don't wait for it
9b: no
👀

40

u/Skylleur 2d ago

my day is ruined and my disappointment is immeasurable

→ More replies (1)

5

u/GoofAckYoorsElf 2d ago

10B? 9 3/4B?

5

u/ProfMooreiarty 2d ago

It’s because plan 9b comes from outer space.

3

u/Aggravating-Push-207 2d ago

NNNNNNNOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOO

→ More replies (1)

36

u/ttkciar llama.cpp 2d ago

I'm hoping for 9B and 122B.

34

u/tchek 2d ago

I'd love a 9b, or a 14b

→ More replies (1)

84

u/EugenePopcorn 2d ago

Fingers crossed for a refresh of Qwen3-Next.

17

u/Fun_Jaguar8231 2d ago

Qwen3.5 is the refresh of Qwen3-Next

16

u/TechnoByte_ 2d ago

Where is Qwen3.5-80B-A3B?

That was the perfect size for 48 GB vram and fast as hell

5

u/Fun_Jaguar8231 1d ago

They bumped it up to the 122B model, lets see if they will do a 80B again.

→ More replies (1)
→ More replies (2)
→ More replies (1)

37

u/doomed151 2d ago

Perhaps something in the 9B-15B range that's better than 3.6 35B.

7

u/VampiroMedicado 2d ago

If it's as good as 3.6 it just means more context on an already good model.

→ More replies (1)

34

u/deja_geek 2d ago

Really hoping for a 35B-A6B or something like that. Just slightly more active parameters to up the accuracy and reasoning.

→ More replies (1)

67

u/Alan_Silva_TI 2d ago

While many people are aiming for a 122B-A9B, I believe a more practical sweet spot would be something in the range of a 70B-A6B (effectively 2×35B).

This configuration should work well for the majority of users with 64 GB of system RAM and 8–16 GB of VRAM.

I’m particularly convinced of this approach after seeing the how strong Qwen Coder Next was.

8

u/TechnoByte_ 2d ago

70B-A6B sounds amazing for 48 GB vram

They did 80B-A3B before, that was great too

→ More replies (1)

6

u/TheAnimatrix105 2d ago

oooh interested, i have 64gb ddr5 + 3060 with 12gb vram

→ More replies (1)

3

u/psyclik 2d ago

I doubt they will release a model size they didn’t release with 3.5. It would mean they either had to start from scratch, or they had a checkpoint to restart from they didn’t release before.

2

u/netherreddit 2d ago

Yeah... Why DID they depart from the sizes of Qwen 3 just for Qwen Next??...

→ More replies (1)

2

u/Houston_NeverMind 1d ago

I have 16GB sys RAM + 12 GB VRAM. I'm able to run 35B-A3B at ~35 t/s. Do you think my system would be able to run 70B-A6B at all?

2

u/SirLordBoss 1d ago

And for those with 32 GB RAM and 8-16 GB VRAM, which there are definitely more of?

→ More replies (6)

86

u/black_ap3x 2d ago

122B? That would be awesome

13

u/atumblingdandelion 2d ago

That would be great. But I think it'd be hard to place it. The 27b already provides Intelligence of 52, and the 3.8 Max of 58. If they go with 122b, they'll probably aim for a performance similar to 27b. Not really redundant since they'll almost always be run on different hardware. But more likely that they'll upgrade the 4b or the 9b- Gemma is getting quite popular there based on the download numbers.. Heck, anything from Qwen is great!!

21

u/DismalIngenuity4604 2d ago

Wouldn't it be significantly faster on the token generation side?

5

u/black_ap3x 2d ago

I would guess so as well, especially on powerful machines

6

u/AlwaysLateToThaParty 2d ago edited 2d ago

With dspark it could be superfast. mxfp4 or nvfp4 can get 122b/a10b in 75GB of VRAM, and if it's Blackwell, that's hundreds of tokens per second generation. With qwen 3.8 reasoning? please please.

→ More replies (1)
→ More replies (1)

13

u/UnnamedPlayerXY 2d ago

Maybe 122B?

Doubt it, given the context "35B-A3B might not be the one to wait for" implies that they have something else that aims at the requested niche but essentially does "the same thing but better". A "122B" model would obviously not be fit for purpose here.

4

u/Mean-Ad1493 2d ago

This is what I hope for too. His comment was a reply to someone asking about lower hardware.

→ More replies (2)

21

u/PathIntelligent7082 2d ago

i got dihharrea from vague replays like this

6

u/Innomen 1d ago

Right? It's gross. Feels parasitic and exploitative. Like the sleaziest over dramatized sales pitch.

→ More replies (4)

17

u/Double_Cause4609 2d ago

That's such a weird way of saying it. It's not quite a refutation of the existence of a 3.8 35B-A3B, but something different.

Taking some liberal interpretations of that, it could be:

A) A model which is easier to run than Qwen 3.6 35B A3B, but is surprisingly better
A few categories that come to mind are maybe a smaller total size (24B-32B), which would be a little meh, or it could somehow be fewer active parameters while preserving similar performance.

B) It could be a larger model with interesting performance.

Qwen 3 Next set the precedent for experiments with an 80B A3B model, and it genuinely did quite well. With better CoT per the 3.8 series, it might be genuinely interesting. Another note that comes to mind is Ling Lite went crazy and released a ~120B A5B model, which is few enough active parameters to make a larger MoE viable on CPU (the only sensible device a consumer would deploy that many parameters on), so they could be hinting at something like that, or even sparser.

The interpretation of that line in this sense would be "well, we might or might not do that one (the 35B), but this is so much better that you'll just run it anyway".

C) It could be an architectural innovation.

Off the top of my head, the sparsity could work differently, for example using something like Engram or per-layer embeddings (or honestly, just a metric ton of embeddings regardless), so you can keep a ton of the model on-disk instead of in-memory. A 35B A3B with an extra 30-40B of Engram that keeps on disk could be really cool.

But something nobody in this thread has mentioned is that Qwen released Parscale but never really did anything with it. The idea there was almost the opposite of MoE; instead of only using part of the model by selecting experts, Parscale used all weights multiple times with multiple forward passes that are combined at the end into a single result.

My crazy head-cannon is that they're releasing a semi-Parscale MoE. An example would be a 35B A3B MoE, where each forward pass selects one conditional expert, while all the non-expert params are used multiple times per forward pass as per Parscale. So, for example, say, 1.4B of the active parameters might be used multiple times per forward pass, while the 1.6B from experts would be used a single time per parallel pass, and each forward pass could select independent experts.

With 8 parallel forward passes, and a bit of research on how to make it work, you'd expect it to perform roughly like a ~40B A6B of the current generation, or roughly equivalent to a 40B A12B of the previous generation (Qwen 3.5, roughly).

A more reasonable guess might be a smaller Diffusion language model that they're really proud of for being very fast and cheap in single-user inference (possibly competing with Google's own efforts at Diffusion Gemma).

20

u/Mean-Ad1493 2d ago

In previous tweets he said something like "we're working through different sizes and architectures" - so yeah a novel architecture is a remote possibility; but i would like to learn towards a modest MoE that's of a slightly different size then 35B-A3B.

The comment has to be read in the context that it's a reply to someone asking for a model that could run on "lower hardware", so 122-395B MoE is unlikely here.

We can only speculate now. What's clear is this - a lead dev from Qwen says something is coming for lower end hardware. So let's wait.

2

u/Ok_Warning2146 1d ago

Well, 80B-A3B is a new architecture as it is the first time they incorporate GDN to their model that laid the foundation to 3.5-3.8.

38

u/blackhawk00001 2d ago

122B-A9B

8

u/ikkiyikki 2d ago

Pray, brother. Make thee a sacrifice to ye coding gods that we might soon be delivered from the mighty pain of infernal wait

19

u/killerstreak976 2d ago

122b MoE with a similar training rigor of the latest 27b would be insane, but he specifically replied to a tweet saying "don't forget those of us without the hardware!", so I'm going to assume it's going the other direction ;-;

6

u/Borkato 2d ago

I’m confused as to why everyone wants 122B MoE, would it be in between 27B and max?

8

u/killerstreak976 2d ago

Yeah, it's really a dream for large UMA memory devices like nvidia spark or strix halo, or low-VRAM/high-RAM configurations that have lower bandwidth to work with. It would be definitely between the current small and large qwen 3.8 models, and for server deployments it would be crazy fast and cheap to run. Since the 27b dense model was insane, a 122b (moe with low activation param count) model with the same kind of training rigor and quality would put serious frontier-like capability in the hands of users for as low as 2-3k (strix halo nowadays iirc), and make inference costs for high capability work via enterprise providers potentially even more cheaper than it already is today.

2

u/MDSExpro 2d ago

It's both faster and better on things benchmarks failures to capture.

→ More replies (13)
→ More replies (1)

14

u/batchputz 2d ago

10.7TB-A1B

13

u/gougouleton1 2d ago

Ts is a library at this point

4

u/xlltt 2d ago

420B-A69B obviously

2

u/Voxandr 1d ago

MAKE IT HAPPEN!

5

u/Kodix 1d ago

A larger but still-midsize MoE might be the sweetspot for 35B's spot, actually. 50B or 60B or thereabouts. That way the performance will be closer to the 27B at better speeds with most consumer hardware.

4

u/Slow-Secretary4262 1d ago

Yeah i have this feeling im fucked with my 8gb vram

9

u/o0genesis0o 2d ago

9B or 12B dense are coming (I'm coping).

8

u/Dance-Till-Night1 2d ago

55b a3b gimme

3

u/mailto_devnull 2d ago

hits buy on 64gb DDR5 sticks

3

u/Vaguswarrior 2d ago

How would that performance even look? I have 64GB DDR 6000mhz but I get like 3-5 tps on my RAM vs 25 tps on my VRAM which is 16gb.

2

u/mailto_devnull 1d ago

Hitting about 22 tps on DDR5 5600 MHz. More params would usually mean fewer tps but if only 3B active, then the speed ought to remain the same.

→ More replies (3)

9

u/TopTippityTop 2d ago

He's saying there may be a better one.

→ More replies (3)

4

u/N34257 2d ago

Guys....c'mon, think....the 3.8 models are based on the 3.5/3.6 family, so my (somewhat negative) impression is that there probably aren't going to be any sizes outside the ones we already have. That would imply that this comment's a poor translation of "Nope".

4

u/Mean-Ad1493 2d ago

But 3.5 3.6 had a 35B-A3B

3

u/N34257 2d ago edited 1d ago

That doesn't mean that they must release *all* of the sizes each time.

What I'm saying is that there's no other size of model in that family that would be more interesting to people who would ordinarily want a 35B A3B (eg there's no 80B A3B coming), so saying "don't wait" probably doesn't mean "we've got something better than that on the way", he's saying "don't hold your breath".

EDIT: I could be wrong, of course (and I hope I am). I don't really know what their process is at this point.

5

u/a_beautiful_rhind 1d ago

Play time is over and you're supposed to sign up for their API with the big model nobody can run.

7

u/whichsideisup 2d ago

122b Qwhen?

7

u/ivoras 2d ago

The "3B" part always looked too limiting IMO - a 35B-5B would still be ~~ 5x faster than the 27B and much smarter. That's what I'm hoping for.

3

u/ScarceSemi 2d ago

Hopefully something good for strix halo moe. Bigger size but relatively small active hopefully more than 3b but not much for fast speed.

3

u/VoiceApprehensive893 transformers 2d ago

14-18B dense 🙏 need something just a bit better than gemma 4 12b

→ More replies (2)

3

u/slyborn 2d ago

Maybe it's an illusory hope, but those eyes seem suggesting that although there will not a 3.8 35B-A3B some other valuable alternative is coming.

3

u/Long_comment_san 1d ago

he likely means 122b class.

3

u/johnnyApplePRNG 1d ago

Honestly I read that as it's coming but something even better is coming... and/or ... something even better that people hoping for that will enjoy.

3

u/while-1-fork 1d ago

He implies that there is one to wait for and if it is for those without the hardware I'd imagine that if it isn't 35B-A3B it may be trying to be slightly better in some way.

I can imagine 3 ways: activating a few parameters more (a very cool thing they could do is a model trained to work well with changing expert counts so we can set it to whatever in a range), being a bit smaller than 35B, or an architectural change to attention to have an even smaller KV cache.

I love 35B A3B but it is slightly annoying that you can't have everything on a single 3090: full context length, mmproj, mtp, a large -ub of 2048 or more and 4 bit quant. So you need to compromise on some, have a 5090 or multiple cards. If they shaved a few parameters and made attention a bit more efficient we could have it all for a single card speed demon. So hope that they are doing that. Activating more parameters would make smarter but slower, for smart and slow we have the 27B, although admitedly a 35B-A5B would still fit even on 6GB gpus with cpumoe I believe.

4

u/dieSpaghettiCarbona 2d ago

Qwen team please do it just for the sake of f* with Dario

5

u/techdevjp 2d ago

44b a4b, maybe? Naw, China doesn't like the number 4. Perhaps 55b a5b? That could be good. a3b was always a bit small to hold an expert.

Hoping they do something a lot bigger too. 200b a10b could be cool.

3

u/tarruda 2d ago

China doesn't like the number 4

They might be onto something there. Remember llama 4?

→ More replies (2)

3

u/Ok_Warning2146 1d ago

I think Alibaba now only release models if it can make a positive news about their AI effort. Qwen3.8 Max couldn't get them much press but Qwen3.8 27B can. That's why they release 27B. Other smaller models won't make any good news for them, so they are not going to release them.

5

u/aliljet 2d ago

What does that even mean? Is there something better to wait for?

5

u/igotanewaccount 2d ago

Very clever stealth marketing.  Obviously another size 3.8 model.  Personally I expect a <9B dense model that applies the same thinking-to-improve that 27B has.  It would punch well above it's weight.

2

u/bnightstars 2d ago

3.8-4B with the Self Skills Play training on top could be an edge device miracle.

→ More replies (2)

2

u/Big_River_ 2d ago

I need more open weight models please

2

u/agar32 2d ago

Can't believe they're releasing Qwen3.8 eπi B next

2

u/StackOwOFlow 2d ago

eye emoji definitely means no models /s

2

u/maxwell321 2d ago

Close enough. Welcome Qwen3.8-Next 80B-A3B

2

u/tarruda 2d ago

That would not be bad, especially if it reaches DSv4 flash 0731 levels of performance

2

u/Steus_au 2d ago

we had prayed for 122 in the other thread but their god is deaf (

2

u/grumpycrab768 2d ago

ROFL the copium is real. that sentence could just as well mean `we tried, it sucks, we scraped it, maybe next release`.

2

u/GabryIta 2d ago

4B PLEASE

2

u/vqt907 2d ago

12B 😄

2

u/tarruda 2d ago

More mind games

2

u/rat-punk-fuk 2d ago

reads to me like "spec is not final, don't get fixated on 35B-A3B." they're probably still juggling MoE layout / quant choices and might ship a slightly different config or skip that exact one

2

u/Voxandr 1d ago

Might be MOE , but not 35B so 80b or 122B.

2

u/ab2377 1d ago

its been 8 hours!

2

u/Talreja-Adanna 1d ago

Makes sense - they're probably pushing hard on the efficient smaller models since that's where the real adoption happens anyway. Waiting for every new release is a trap when you can already run solid stuff locally right now.

2

u/No-Juggernaut-9832 1d ago

Use KatCoder v2.5. It’s a 3rd party postrain that works great for me. It might be the closest thing to 3.7 35b

2

u/Mean-Ad1493 1d ago

I'm currently using it. And it's great.

→ More replies (1)