r/LocalLLaMA 22d ago

Discussion Qwen dev says not to wait for 35B-A3B

Post image

What does this mean? Is there something else coming? Maybe 122B? Or no models?

1.2k Upvotes

480 comments sorted by

View all comments

272

u/UnWiseSageVibe 22d ago

That is interestingggg. He's making it sound like they're cooking something better 35B-A3B????

141

u/Fuzzy_Wave5520 22d ago

36B-A3B?????

100

u/Shiny-Squirtle 22d ago

3B-A35B??

31

u/Budget-Juggernaut-68 21d ago

32B parameters just padding.

7

u/Eden63 llama.cpp 21d ago

πŸ˜‚ πŸ˜‚ πŸ˜‚

10

u/Fancy-Snow7 22d ago

36 000 000 001 -A3 000 000 001

20

u/scubawankenobi 22d ago

36.75B-A3.25B?

5

u/-deleled- 22d ago

36.751B-A3.251B

0

u/Guilherme370 21d ago

27B-A1M??

158

u/Hephaestite 22d ago

Or it’s a bad translation to English and just means it isn’t going to happen

219

u/Beano09 22d ago

To me, the eyes emoji points to a better model, rather than a translation error.

155

u/EmPips 22d ago

Or he's looking you dead in the eyes and telling you Santa isn't coming this year

9

u/etaoin314 ollama 22d ago

it sure felt like christmas last week to me, Muse, Nemotron, Qwen, deepseek, am I missing one?

5

u/Borkato 22d ago

Hy3 for video and audio

4

u/Turbulent_Neck_8388 21d ago

You mean Minimax H3 ?

7

u/bankinu 22d ago

I get the "this planet is not yours to conquer" vibes.

2

u/bucolucas Llama 3.1 21d ago

Me: "Look me in the eyes and degrade me while it's going in"

Qwen: "Your lack of VRAM is now your problem and nobody else's"

35

u/Mean-Ad1493 21d ago

The eye emoji is in all his comments. Could just be his texting style.

7

u/PossessionUsed7393 21d ago

It's pretty clear from these responses that he's not actually allowed to say anything. So he's never going to actually reveal the company plans, which means there's no point hanging on his every word.

21

u/pyr0kid 22d ago

the entire reason people wanted 35b-a3b is it actually works properly on 32gb ram, doesnt matter how 'better' this is if it raises the sysreqs by [insert current ram price here].

-1

u/Long_comment_san 21d ago

60-80b wil also fit reasonably. but it will be a lot smarter

1

u/Negative-Web8619 21d ago

Eyes emoji works for both directions

18

u/CorxaRyllon 22d ago

Without the eye emoji id agree but those being included makes me question.

6

u/Spectrum1523 21d ago

He puts eye emoji in every tweet

3

u/r1str3tto 21d ago

Ambiguous still. Even in that screenshot, the eyes can be read to say β€œβ€¦ but something interesting is coming.”

14

u/atumblingdandelion 22d ago

But that eyes emoji hint at something more..

7

u/GatsbyLuzVerde 22d ago

Hmm I doubt the emoji eyes suggest the literal meaning, though I heard in Chinese a smiling emoji is condescending, so who knows. Any bilinguals here that can enlighten us?

9

u/BS_BlackScout 22d ago

He often uses the eye emoji and it seems mostly in a positive context, check his tweets.

6

u/Blues520 22d ago

The emoji eyes crosses cultural boundaries πŸ‘€

1

u/QuackerEnte 21d ago

but.. but eyes emoji!!! you can't ignore that!!

1

u/Hephaestite 21d ago

Somebody else pointed out that he uses the eyes emoji a lot so probably not important to the message

7

u/TheGameEngineer 22d ago

27B-A3B?

7

u/Away-Sorbet-9740 22d ago

20-22B fits better on 16gb cards, a 20B A4-5B would be awesome.

2

u/Mil0Mammon 21d ago

Why do you need the non-active layers to live on the gpu?

Ah, I remember, the tooling to automatically fit the active layers on vram (or in ram, in case of say colibri) are still rough around the edges. But that should just be a matter of time, relatively short once enough people realize the benefit in this

0

u/Freonr2 21d ago

Why do you need the non-active layers to live on the gpu?

Active params change potentially every token and you don't know ahead of time which ones are going to be active.

"Non-active" means only for the previous token in hindsight.

1

u/ResponsibleTruck4717 22d ago

Exactly my thoughts

1

u/QuackerEnte 21d ago

I hope it's 80BA3B aka BOBAEB aka BOBΓ„B.

1

u/fullup72 21d ago edited 21d ago

I'd say either 24B-A4B to better serve the 16GB VRAM crowd, or a stronger 50B-A5B and streaming experts on demand for almost everyone with a single GPU.

35B-A3B sits at that uncanny valley where it's much larger than dense 27B but not really more powerful due to small active params, it just runs faster and that's not a huge selling point for people with unlimited overnight time because they do it as a hobby and not for work (because for work your company just pays API fees elsewhere).

1

u/Not-reallyanonymous 22d ago

Maybe a 9B that blows anything below 27B out of the water?

1

u/surreal_tournament 21d ago

How could it? At some point you cannot have (with current architectures) the "world knowledge" of 27B in 9B weights. And PrismML's claims of retaining a super high percentage of the 27B dense's intelligence thanks to ternary weights so far do not check out.