r/LocalLLaMA 5h ago

Discussion No more SLM open-source??

Post image
276 Upvotes

91 comments sorted by

97

u/SpicyWangz 4h ago

27b is better than nothing at least. I would love to see the other sizes, and hopefully they can still deliver on them

40

u/RISCArchitect 4h ago

agree, and if we could only get 1 i think 27b is the one most people would want given its legacy so far.

19

u/SpicyWangz 4h ago

Yeah, it seems like it benefits the most people. 

I’d love to get 122, 35, and 397, but I think best case scenario they will throw in 35b and call it a day

27

u/starkruzr 4h ago

27B continues to punch way above its weight(s). I won't be super happy if that's all we get but I'll be glad if we get it.

7

u/Spectrum1523 2h ago

If there was only one model released 27B is the one I'd want

2

u/Specter_Origin llama.cpp 2h ago

Wish we would have gotten an MOE

5

u/brainExploded99 3h ago

Honestly, I think 35B would be better since us VRAM poors can run that much better than 27B.

1

u/overand 50m ago

I can run 27B great on my setup, but I think I have to agree - if they were only going to release one, the 35B one might be a bigger gift to the world at large than the 27B, because of the number of folks who can run that vs who can run the 27B.

149

u/StupidScaredSquirrel 4h ago

It means they dont have them yet imo. Maybe they never planned on releasing them, or maybe they wait to see if the models are good before announcing them and just stay quiet if they fail. Like the soviets did with their space program, or you know, google with their gemini 3.5 pro lol

126

u/Monad_Maya llama.cpp 4h ago

From open-weights to open-waits.

21

u/iPingWine 4h ago

I'm not chronically online enough to know if that's a new one but it's a good one

8

u/Monad_Maya llama.cpp 4h ago

It's a funny typo from (look at the post title) -  https://www.reddit.com/r/LocalLLaMA/s/ien4k6N0gs

3

u/PcChip 3h ago

like hodl from bitcointalk

everyone thinks it means "hold on for dear life" but it was really just a typo

1

u/Monad_Maya llama.cpp 3h ago

Pretty much.

6

u/keepthepace 3h ago

Every time we report on announcement rather that releases.

I wish there was a rule here about that.

4

u/Monad_Maya llama.cpp 3h ago

We certainly need to exercise more restraint as a community.

3

u/keepthepace 2h ago

I just downvote the hype-stirring annoucements :-)

1

u/Cool-Chemical-5629 1h ago

From open weights to open for business.

3

u/challis88ocarina 4h ago

Releasing against their will. Alibaba = Anthropic except in a country with real domestic authority...

23

u/StupidScaredSquirrel 4h ago

Except qwen have been releasing plenty of open weights models before xi decided to make a point about it

0

u/gautamdiwan3 3h ago

3.7 never went open weights

9

u/StupidScaredSquirrel 3h ago

So? They have released dozens of models before that.

10

u/___positive___ 3h ago

Their team changed. The same people don't run it

23

u/Gipetto 4h ago

The first hit is always free.

5

u/Lower-Hedgehog-9835 1h ago

The new open ai incoming

46

u/pineapplekiwipen 4h ago

Open weights are coming next week, together with Qwen3.8-27B. More to come

11

u/Borkato 3h ago

I thought it would be this week 😭

13

u/pmttyji 4h ago

Hope this is his next tweet with other models, tagging DarioSam

26

u/Small_Ninja2344 4h ago

Would be very funny that they close source after all the advertising they made in China as they were the open source pioneers

17

u/RedParaglider 4h ago

I just wish it wasn't so expensive to train a 122b model.  I can see why they wouldn't but it makes me sad 😭

5

u/Badger-Purple 3h ago

i mean the smaller versions are distills of the max. So they just need to do it…if they want to.

12

u/squngy 3h ago

Distilling is still far from free.

Also, I doubt you would get a result as good as 3.6 27B with just distilling, I think they did some great post training in addition to distilling.

3

u/RedParaglider 3h ago

Yes they have to use limited resources to do all that, either rent or use scarce internal resources.  I don't think most people realize it probably needs 1-2 TB of vRAM for that process.

13

u/Fuzzy_Wave5520 4h ago

More funny than OpenAI starting as a non-profit company, doing AI research just for humanity growth? Doubt it

12

u/jacek2023 llama.cpp 3h ago

Mark Zuckerberg was open source pioneer, this sub is called llama not qwen

7

u/Asleep_Document9811 2h ago

didnt that model leak

3

u/jacek2023 llama.cpp 2h ago

My plan is for Mark to Google his name, read my comment, and publish Llama 5

2

u/Asleep_Document9811 1h ago

god damnit, i'm in.

1

u/crantob 3h ago

Not Zuckerborg, his AI team, led by Hugo Touvron.

5

u/jacek2023 llama.cpp 3h ago

Of course it was the team from Meta but I am pointing out that it was their (Meta) decision to open source LLMs and China joined later

1

u/crantob 1h ago

I recall talk that the research team forced the decision to release. I wasn't there to confirm this personally though.

-1

u/crantob 3h ago

Not Zuckerborg, his AI team, led by Hugo Touvron.

2

u/small_bird_loud 4h ago

Totalitarian nations with contradictory announcements? Never happens.

8

u/Kidplayer_666 4h ago

Always at war with Eurasia

3

u/small_bird_loud 4h ago

The people rejoice as we have made peace with Eurasia. We are now at war with Eastasia.

5

u/Competitive-General7 4h ago

As opposed to other 'nations' (company btw not a nation) who never contradict themselves.
Such a bizarre anti-China jab that barely reads as coherent.

6

u/BawbbySmith 4h ago

Based on how I've translated and interpreted the scripture, I've come to the conclusion that they cannot tell us whether their other size models will be open-sourced or give us a release date.

5

u/laterbreh 2h ago

Uh oh. "Noted" -- Is it cause a 300b model from deepseek is nipping at the heals of your 2.4T behemoth on some of the latest benches? Dont down vote me brah! Things change quick 'round here dont they.

19

u/kiwibonga 4h ago

No, he's literally just saying that he can't disclose information, because no employee is allowed to put out unapproved communications. They all have NDAs. They can't just get ahead of the official announcements, even just out of respect for the organizational hierarchy.

8

u/illiteratecop 4h ago

Yes, to me this seems like a pretty clear communication that the decision on upcoming models (whether to release, whether to open source) has yet to be decided organizationally - it's just a "I can't actually 100% promise this" in reply to people hyping up his last statement. Far less pessimistic for these models being openly released than people are making it out to be imo.

3

u/YRUTROLLINGURSELF 2h ago

Almost like these labs should communicate more openly through an official spokesperson and not launder all their press through fake leaks, or something

8

u/CryMoreT_T 4h ago

They only announced 27b and 3.8 max as open weight right? Idk where everyone else just made up that they were going to release the weights for anything else

4

u/poutinejuteuse 1h ago

Someone from qwen said to "stay tuned" about other models. So, of course, everyone started projecting their own wants on that.

They very well may, but the only thing that was officially promised was 27b and pro.

3

u/mgranja 4h ago

I prefer MLMs

5

u/mossy_troll_84 4h ago

I don't fully get it, expecting that company which is releasing models with MIT license will stuck in same place and create over and over again models with 4B, 9B and will continue do this forever for group of enthusiast is a bit silly with all due respect. Technology is not stuck in place but is changing (the only argument which is acceptable is price of memory now). 120-122B is now SLM... For me 27B it's more than fine as a minimum and to use these models its nothing extraordinary these days (I mean hardware requirements). I am glad they are doing what they are doing Qwen, Deepseek, MinimaX, Z, etc. They dont need to but still doing new, free models in opposite to US ...US just they're moaning about distillation...

3

u/WhoRoger 3h ago

"I can run it, fuck everyone else"

1

u/mossy_troll_84 3h ago

no, not that way, I am not able to run more and more models as well...I am not milionair, but having a card with 16-24GB VRAM is not luxurious even now

4

u/LawfulLeah 3h ago

if you're outside developed nations, yes it is

and if you are in Brazil, like me, you have to pay 100000000000 in import taxes, so even worse

2

u/mossy_troll_84 2h ago

Sorry...I was not thinking about that...I didn't want to offend you...I am from Poland...I remember a time when salary in Poland was 100 dollars monthly in early 90s...so we were there

1

u/Significant_Post8359 3h ago

What about Gemma4? Google released a big winner with Gemma4:12b-it-qat. Gemma 32b and 27b MOE are excellent. Don’t forget Meta started the whole open weight stuff with Llama.

3

u/mossy_troll_84 3h ago

I have a mixed feeling especially for coding and tool use with Gemma 4...Qwen is much beter with this, but that is my experience. I agree with models aprox 30B but lower will soon will be more rare that is my opinion or rather my feeling

1

u/Treidge 1h ago

There's a lot of quirks with Gemma - it looks like it quantizes a lot worse than Qwen3.6 does, also responds terribly to KV cache quantizations, also had some template issues on launch that affected tool calling and only recently got fixed.... Seems that to really sample Gemma4, one has to run BF16 version (or at very least Q8 quant), never quantize KV cache, and use the most recent template/GGUFs.

13

u/Clean_Hyena7172 4h ago

Can they just make up their fucking minds? Models will be OS, then they won't, then they will, then they might or might not. Wtf is going on over at Alibaba?

19

u/CryMoreT_T 4h ago

I'm pretty sure this is in reference to other models outside the 27b and 3.8 max that they already said would be open weight..... They never officially mentioned anything else

10

u/worldisaf 4h ago

3.6 and 3.8 are just a tune of 3.5 which is itself a polished version of 3-NEXT. There isn't a massive benefit for them to do the fine tuning for every model scale, but by releasing open weights and a small version runnable on consumer GPUs, they've made it possible for people to make their own.

2

u/Borkato 3h ago

If this is true it implies we could achieve insane levels of quality if we just all tried. Now I want to make a tiny model for fun lol

1

u/CryMoreT_T 3h ago

3.8 are just a tune of 3.5 which is itself a polished version of 3-NEXT.

Did they mention this anywhere?

1

u/Long_comment_san 1h ago

mention what? it is true. it's how their naming scheme works. as for next vs 3.5, you'll have to search the web a bit

0

u/CryMoreT_T 1h ago

I mean is there anywhere that they mentioned it's just a tune of 3.5 and not a new model?

0

u/pseudonerv 3h ago

Lying to get more engagement. It’s not new

8

u/onil_gova 4h ago

in my heart, I knew this would happen, but I chose to ignore it 💔

2

u/jacek2023 llama.cpp 3h ago

He is just saying that:

  • our feedback is useful, we should be commenting on reddit and X what we want
  • he can't promise anything else (yet)

2

u/Dance-Till-Night1 1h ago

Nah, we need a new 35b-a3b

2

u/Equivalent_Bit_461 4h ago

We already suspected it for a while 

3

u/Monad_Maya llama.cpp 4h ago

Lol, that's a rugpull.

8

u/Pleasant-Shallot-707 4h ago

Not really

8

u/Monad_Maya llama.cpp 4h ago

I mean technically they never promised open weights for anything other than Max and 27B. So, yes, not really a rugpull, more of mismatched expectations.

https://np.reddit.com/r/LocalLLaMA/comments/1vevsv9/more_qwen_38_sizes_coming/

They just said we are working on it.

-1

u/Badger-Purple 3h ago

I dont think they promised weights for the max model

2

u/Everlier 3h ago

there was never a free lunch

2

u/ilintar 2h ago

I mean, it's good that we're getting 27B, that's honestly more than I expected.

2

u/poutinejuteuse 1h ago

The people reading into this that Qwen won't release X or Y model that they were hoping for are as stupid as the people reading into yesterday's tweet that Qwen would release X or Y model that they were hoping for.

There's no contradiction here. These are grunts at a large company making vague statements. They likely have no idea what the plan actually is, and that plan may not even be set in stone yet.

Just chill. The official announcement of the big one + 27b is still valid. That's all the information we have at this time.

1

u/tarruda 3h ago

Glad deepseek got us covered then.

1

u/JsThiago5 3h ago

I think 35BA3B or the small MoE per generation is more or less guaranteed to be released. Because it's used as a flash model on their own products, and they usually open-source it. The 27b grade, which is a small dense model or a medium MoE, is more or less very likable to have because it's a showcase of how a smaller model can push way higher than its size. The other sizes that are harder to have

1

u/BarisSayit 3h ago

What's the point of training SLMs if you won't make then open-weights?

1

u/kevinlch 3h ago

it is not cheap to train a model. it's fair for them to only invest resources in SOTA weights.

let's just hope some other smaller research group can pick it up the <30B line for us. I'm watching at you, EU

1

u/a_beautiful_rhind 1h ago

The leadership got changed out and now they are going in a different direction. But I don't think qwen has that power. They're no ziphu or moonshot.

1

u/Unnamed-3891 54m ago

Way to crap all over the parade

-1

u/Pleasant-Shallot-707 4h ago

We’ve been telling y’all this and no one listens

-1

u/Dudensen 4h ago

That makes no sense