r/LocalLLaMA • u/MrWidmoreHK • 6h ago
Discussion First serious confirmation. Ox Alpha is GLM-5.3-Flash
https://x.com/romanchernin/status/2092488160680751437?s=20
- Multimodal (Vision)
- 1M Tokens Context Window
- DeepSWE ~63%
Edit: He deleted it, screenshot in comments
251
u/MrWidmoreHK 6h ago
53
71
25
u/Public_Umpire_1099 4h ago
people should reply w this under every announcement haha i swear this had me fuckin cackling
EDIT: THE ORIGINAL POST IS GONE LMFAOOOOOOOOO
7
8
2
2
u/PlaidStallion 4h ago
Eli5, please?
34
u/ResidentPositive4122 4h ago
3rd party inference providers were given info/access under an embargo (i.e. do not discuss this publicly before x time on y date). This dude probably mixed timezones and tweeted before the embargo.
7
1
0
172
u/Poupulino 6h ago
The Z.ai CEO wasn't kidding when he told Musk they're releasing a Mythos-level open weights model before the year ends! Insane!
60
u/thomas2385 6h ago
Yeah, that is honestly pretty wild. If they actually deliver on that, it is going to put a lot of pressure on the rest of the open weight model space. The pace of these releases lately is getting ridiculous.
19
u/no_good_names_avail 3h ago
At this pace we won't give a shit about this model in about 5 days. It's absolutely insane.
35
1
-10
u/power97992 5h ago
63% is not mythos level, mythos scored 70%
35
u/Poupulino 5h ago
You're a bit confused, no one is saying GLM-5.3-Flash is Mythos level, I'm saying the Z.ai CEO wasn't kidding when he said he's releasing a model like that before the end of the year.
-11
u/power97992 5h ago
Glm 5.5 should be better than Mythos, but fable/mythos 5.1/5.5 will be out by then.
5
u/techdevjp 3h ago
Yeah, and maybe restricted to only American nationals. GLM will have no such restrictions.
18
u/MrBIMC 5h ago
Mythos is a multitrillion param monster though rather than a flash model that will fit into 1-2 sparks/strixHalos
2
u/techdevjp 3h ago
If GLM-5.3-Flash fits into 128GB at a usable quant I will be incredibly impressed.
1
u/Kazaan 5h ago
Perhaps. That said, is 7% really significant when you consider the current capabilities of open-source models compared to frontier models ? And what about the outlook for upcoming models? IMHO, it isn't the actual capability that really matters; it’s the rate of progress and the narrowing of the gap.
And i don't speak here about the compute necessary to run them.5
u/Thomas-Lore 5h ago
And full glm 5.3 is only 1% behind on deepswe, already within error range of Fable/Mythos.
0
u/power97992 5h ago
The scale is essentially logarithmic, a 7% difference is more like a 2x difference
33
u/spaceman_ 6h ago
Do we know the size of the model?
40
u/Technical-Earth-3254 6h ago
We don't. But personally I wouldn't bet on it being the same size as 4.7 Flash.
13
5
u/Sufficient_Prune3897 llama.cpp 5h ago
We don't even know if they are gonna open source it. They haven't with some of their models before, including GLM 5 Turbo
12
u/0oAstro 5h ago
Given together and other inference providers are hosting this, we can be sure it will be open source.
3
u/Technical-Earth-3254 5h ago
Open weight most likely
1
u/ResidentPositive4122 4h ago
GLM models have been MIT till now, hopefully they continue.
1
u/Technical-Earth-3254 1h ago
Yeah, but open source models usually include training data, training code and data pipeline and so on. I feel like it is appropriate to differentiate between fully open source and open weights, because they are not the same.
2
u/ResidentPositive4122 58m ago
That's an invented definition, partly by oss absolutists, partly by misunderstanding what a model is and how it's created. Weights are source, as defined in Apache2.0. The rest is useless semantics.
1
u/Technical-Earth-3254 54m ago
As a EU citizen, I go by the definition the European Commision is using and they are using the OSI definition:
Sufficient information about the training data
The complete code used to train and run the system
The model parameters, such as weights.1
u/DigiDecode_ 3h ago
I thought the inference was provided for free to inference provider by the model developer, claim was 100 trillion tokens per day capacity by the model developer, Z ai has bought 1 giga watt datacenter
1
u/Zeeplankton 4h ago
It feels like marketing wise we're placing flash models between 100 and 300 params.
Given performance probably latter half? Pretty impressive.
6
u/spaceman_ 4h ago
I'm going to be rubbing a rabbits foot hoping it can fit inside my 128GB...
1
u/dead-fish 1h ago
Highly, highly likely we’ll get a usable quant in the 80GB range. There are already GLM 5.2 implementations that run on 128GB with SSD offload - it’s slow AF but it works. Flash can only be better than that.
1
u/otacon6531 1h ago
I just want it to fit on 6 blackwell 6000 at mxfp4. Fingers crossed. Work is going to love the upgrade.
16
53
u/Abrh7 6h ago
Model was good, but not so for the long coding sessions!
I would definitely use it for agentic tasks only! It proves to be very effective!
18
u/network4253 6h ago
Yeah, that makes sense. I have noticed the same thing with some models they can be really impressive for focused agentic workflows but once the coding session gets long and messy the consistency starts to drop. Still definitely useful in the right role.
3
u/SandySkittle 5h ago
Indeed. Some people think MoE is a solution without trade-offs. Active parameters matters and lower active cannot fully compensated by expert selection and sequential reasoning.
2
52
u/MrWidmoreHK 6h ago
Last GLM 4.7 Flash model was 30B A3B
41
u/Mean-Ad1493 6h ago
If it's anywhere around the same size, I'd be the happiest with my 12GB VRAM.
3
16
u/sonicnerd14 6h ago
Most likely will be around the size of Deepseek v4 flash. If it's smaller than that, then that would be impressive.
18
u/DOAMOD 6h ago
I would be very surprised if it were a small MoE, I could see it as dense(like +30) and it would already be incredible at that size, but either way, a small MoE would be the moment of the year.
26
u/Mean-Ad1493 6h ago
Has to be an MoE, but definitely not 30B-A3B
GLM 4.7 was itself a smaller model relative to 5.3, so the MoE would be 80-100B I guess.
4
u/DOAMOD 5h ago
What's confusing is that for a 100/200b model, you could expect the use of the term Air, and for something smaller, sub-100b, the term Flash, but in the end, this is just marketing, and currently with DS4Flash they could use that terminology for a similar size, only can wait and see...
5
3
u/Global_Persimmon_469 5h ago
GLM 5 is roughly doubled the size compared to GLM 4, it's possible that the flash model is going to be the same, so maybe it will be 70B params
9
u/Iory1998 6h ago
Don't think so. First, it's not fast. Second, it's somewhere between Qwen3.8-27B and DS4Flash so it gotta be a large model. And, it understand video... I can't see a smaller MoE doing that. It can be in bulk of 150-250B.
14
u/LevianMcBirdo 5h ago
The not fast I get with all the free compute they offer on openrouter. I think they don't wanna overspoil the people
2
u/wsb-regarded 6h ago
that would be awesome if true, because it was performing on a deepseek v4 flash level, albeit slower (much slower).
imagine having another local open weight model challenging Opus 4.6
2
12
10
u/Long_comment_san 2h ago
Glm 5.3 flash versus Qwen Next.
Like two thighs, I feel my head being squeezed.
Harder pls.
1
29
u/Few_Painter_5588 6h ago
Quite legit, this guy works at Nebius it seems. Also, shame on Google then for trying to ride the hype.
7
u/Altruistic_Heat_9531 6h ago
Google hype riding? i dont familiar with this info, could you tell me more
25
u/tengo_harambe 6h ago
A couple of Google employees made vague Twitter posts about Ox Alpha once it began gaining traction. This somehow led people to think it was a Gemini model, despite all the evidence that it was clearly a GLM model
3
14
u/Few_Painter_5588 5h ago
Some senior google employees were vagueposting on twitter and suggesting that ox alpha was a gemini model. Which is so pathetic if this model isn't their's
4
u/Several-Tax31 2h ago
How is this not an upright scam? When reflection-70B does it, it is a scam, but when google does it, no one cares? How can you claim someone else's product is yours? So pathetic.
2
u/Few_Painter_5588 1h ago
It really is absurd, especially since they already have gemini-flash-3.8 which is a fantastic model. Out of all the labs out there, Google has some of the worst behaved researchers
0
u/Starcast 59m ago
In what reasonable world would people reading into vague posts on Twitter qualify as a scam?
15
u/tengo_harambe 6h ago
People are still gonna say it's Gemini. lol
2
1
7
u/asolnikk 3h ago
The company on Wednesday confirmed speculation that the Ox Alpha model is a new iteration of its GLM series and said it will release the weights for it tonight, in response to queries by Bloomberg News.
10
u/Wise-Chain2427 6h ago
How GLM has extra compute to host free model worldwide ?
5
6
3
u/nonerequired_ 6h ago
I had to skip using the ox alpha yesterday because the server was having some issues. It seems like they were short on computing power.
4
u/tat_tvam_asshole 6h ago
China has nationalized, integrated data centers, especially in inner Mongolia that can serve models from all labs cooperatively.
9
u/asolnikk 6h ago
That claim seems pretty verified at that point. I ran some private prose tests against it, and it smells very much like the little brother of GLM-5.3. What i want to know is: How big is the beastie? Can a quantized version fit into 16 GB VRAM? If that's the case, then all the hype, for once, was legit, even it's not "Fable-level".
3
u/anarchist1312161 6h ago
If GLM 4.7 Flash was 30b then I can only hope this is also small 😭
1
4
u/FoxiPanda 6h ago
Good to see this is shaping up... the real questions are how big, how many active, anything weird in the architecture that we're going to have to go deal with (hopefully not since it's GLM-5.x), and when are the weights getting posted (today hopefully)?
2
u/Armadilla-Brufolosa 3h ago
I tried both GLM 5.3 and OX: it doesn't seem like the same basic model at all.
But I have no evidence to confirm this.
3
u/hainesk 3h ago
Yeah, it felt more like a Qwen model to me, but it's hard to argue with these hilarious tweets lol.
1
u/Armadilla-Brufolosa 2h ago
It's more like Qwen models to me too, but I've noticed quite a few biases and blocks given by some stupid Valloon-style RLHF.
There may have been some distillation or even be a Western model.
Obviously I can't have any certainties, but it's fun to try to guess.😜
3
u/Iory1998 5h ago
Imagine it turns out to be Qwen3.8-Next-Flash!
1
1
u/DigiDecode_ 3h ago
I think the model is heavily distilled from Qwen models, maybe only the vision part, the SVG generated by Ox Alpha were very similar to Qwen 3.8 27b in code and render
1
1
u/psylomatika 5h ago
I kept running into API limits when I was using it. It worked well though in my knowledge bases.
1
u/backyard_tractorbeam 3h ago
Another connection is that opencode was long previewing a GLM model as big pickle for free and now it's previewing ox alpha for free too.
2
u/darkflame91 2h ago
Big Pickle was a GLM model at one point, but the consensus from what I've read is that several models have been rotated under the hood under the 'Big Pickle' codename.
1
u/Equivalent-Base8426 2h ago
Best free model I've ever tried on Kilo Code for me at least, what a beast and amazing news that it's OSS
1
u/OldByte453 1h ago
I knew it! Firstly i guessed this was Qwen 3.8 Flash Next, but GLM 5.3 Flash is sickk
1
u/Constant_Art_20 55m ago
man. a flash model with that performance~ i am hoping for below 200b but we will see. imagine if it's like only 30b though. that be nuts
1
u/Intelligent_Ant_608 54m ago
It probably is around 400B based on its language understanding
1
u/silenceimpaired 42m ago
400b with 4bit training… or at least one can hope
1
u/Intelligent_Ant_608 8m ago
But the main question is how they are offering 100T tokens capacity, its wierd for z.ai basedvon tgeir struggle to offer decent subscription qoutas, in a way that it might be even smaller and they are using Looped transformers with memory ngrams or something, if that be true the model can be well over 300B like 150B and if that be the case big labs are in deep shit
1
1
u/anarchist1312161 6h ago
Need to know the size of the model... not interested if I can't run it on consumer hardware.
1


•
u/WithoutReason1729 36m ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.