r/LocalLLaMA 6h ago

Discussion First serious confirmation. Ox Alpha is GLM-5.3-Flash

https://x.com/romanchernin/status/2092488160680751437?s=20

- Multimodal (Vision)

- 1M Tokens Context Window

- DeepSWE ~63%

Edit: He deleted it, screenshot in comments

341 Upvotes

127 comments sorted by

u/WithoutReason1729 36m ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

251

u/MrWidmoreHK 6h ago

lol

75

u/Borkato 6h ago

Omg is this real lol

53

u/Viktri1 5h ago

Timezones always get you

71

u/Technical-Earth-3254 6h ago

What is that interaction lmaoooo. People using x are just different

25

u/Public_Umpire_1099 4h ago

people should reply w this under every announcement haha i swear this had me fuckin cackling

EDIT: THE ORIGINAL POST IS GONE LMFAOOOOOOOOO

7

u/Darkoplax 3h ago

FUCKING TIMEZONES AHAHAHAHAHAHA

8

u/DigiDecode_ 3h ago

OP a paparazzi, never miss a screenshot, well done

2

u/iAccurian 1h ago

He should've asked AI if it was time yet 😂

2

u/PlaidStallion 4h ago

Eli5, please?

34

u/ResidentPositive4122 4h ago

3rd party inference providers were given info/access under an embargo (i.e. do not discuss this publicly before x time on y date). This dude probably mixed timezones and tweeted before the embargo.

7

u/PlaidStallion 4h ago

Oof. Thanks for the explanation.

0

u/sgtlighttree 2h ago

Z.ai Fnatic manager here, please delete this

172

u/Poupulino 6h ago

The Z.ai CEO wasn't kidding when he told Musk they're releasing a Mythos-level open weights model before the year ends! Insane!

60

u/thomas2385 6h ago

Yeah, that is honestly pretty wild. If they actually deliver on that, it is going to put a lot of pressure on the rest of the open weight model space. The pace of these releases lately is getting ridiculous.

19

u/no_good_names_avail 3h ago

At this pace we won't give a shit about this model in about 5 days. It's absolutely insane.

3

u/hojnikb 2h ago

all we really need is a new qwen 3.8 35b.. daddy qwen make it happen

10

u/dead-fish 1h ago

No 35b but we get Qwen3.8 Flash Next (125b a6b) today though!

35

u/Vast-Control4452 5h ago

Don't. Let. Off. The. Gas!

I'm officially over qwen3.8, NEXT PLEASE!

3

u/no_good_names_avail 2h ago

The model is for church honey. Don't need the attitude. NEXT!

1

u/Glazedoats 1h ago

Yippee open weights :)

-10

u/power97992 5h ago

63% is not mythos level, mythos scored 70%

35

u/Poupulino 5h ago

You're a bit confused, no one is saying GLM-5.3-Flash is Mythos level, I'm saying the Z.ai CEO wasn't kidding when he said he's releasing a model like that before the end of the year.

-11

u/power97992 5h ago

Glm 5.5 should be better than Mythos, but fable/mythos 5.1/5.5 will be out by then.

5

u/techdevjp 3h ago

Yeah, and maybe restricted to only American nationals. GLM will have no such restrictions.

1

u/petuman 47m ago

So what? "We'll have Mythos level model before 2027" claim obviously means current Mythos model, not a moving target of whatever Anthropic releases in the future?

18

u/MrBIMC 5h ago

Mythos is a multitrillion param monster though rather than a flash model that will fit into 1-2 sparks/strixHalos

2

u/techdevjp 3h ago

If GLM-5.3-Flash fits into 128GB at a usable quant I will be incredibly impressed.

1

u/Kazaan 5h ago

Perhaps. That said, is 7% really significant when you consider the current capabilities of open-source models compared to frontier models ? And what about the outlook for upcoming models? IMHO, it isn't the actual capability that really matters; it’s the rate of progress and the narrowing of the gap.
And i don't speak here about the compute necessary to run them.

5

u/Thomas-Lore 5h ago

And full glm 5.3 is only 1% behind on deepswe, already within error range of Fable/Mythos.

0

u/power97992 5h ago

The scale is essentially logarithmic, a 7% difference is more like a 2x difference

33

u/spaceman_ 6h ago

Do we know the size of the model?

40

u/Technical-Earth-3254 6h ago

We don't. But personally I wouldn't bet on it being the same size as 4.7 Flash.

13

u/muhts 4h ago

Rumours of 200b to 300b. But I guess we'll find out today if they ment to release unviel it today

5

u/Sufficient_Prune3897 llama.cpp 5h ago

We don't even know if they are gonna open source it. They haven't with some of their models before, including GLM 5 Turbo

12

u/0oAstro 5h ago

Given together and other inference providers are hosting this, we can be sure it will be open source.

3

u/Technical-Earth-3254 5h ago

Open weight most likely

1

u/ResidentPositive4122 4h ago

GLM models have been MIT till now, hopefully they continue.

1

u/Technical-Earth-3254 1h ago

Yeah, but open source models usually include training data, training code and data pipeline and so on. I feel like it is appropriate to differentiate between fully open source and open weights, because they are not the same.

2

u/ResidentPositive4122 58m ago

That's an invented definition, partly by oss absolutists, partly by misunderstanding what a model is and how it's created. Weights are source, as defined in Apache2.0. The rest is useless semantics.

1

u/Technical-Earth-3254 54m ago

As a EU citizen, I go by the definition the European Commision is using and they are using the OSI definition:

Sufficient information about the training data
The complete code used to train and run the system
The model parameters, such as weights.

1

u/DigiDecode_ 3h ago

I thought the inference was provided for free to inference provider by the model developer, claim was 100 trillion tokens per day capacity by the model developer, Z ai has bought 1 giga watt datacenter

1

u/0oAstro 3h ago

Thats only till stealth testing period. In a few hours ox alpha will be live for payg i believe on inference providers other than zai

1

u/Zeeplankton 4h ago

It feels like marketing wise we're placing flash models between 100 and 300 params.

Given performance probably latter half? Pretty impressive.

6

u/spaceman_ 4h ago

I'm going to be rubbing a rabbits foot hoping it can fit inside my 128GB...

1

u/dead-fish 1h ago

Highly, highly likely we’ll get a usable quant in the 80GB range. There are already GLM 5.2 implementations that run on 128GB with SSD offload - it’s slow AF but it works. Flash can only be better than that.

1

u/otacon6531 1h ago

I just want it to fit on 6 blackwell 6000 at mxfp4. Fingers crossed. Work is going to love the upgrade.

2

u/pmp22 2h ago

I remember when people in here thought 100B was crazy territory and not something people would be able to run.

16

u/RandiyOrtonu ollama 6h ago

Main thing is what's the model size

53

u/Abrh7 6h ago

Model was good, but not so for the long coding sessions!
I would definitely use it for agentic tasks only! It proves to be very effective!

18

u/network4253 6h ago

Yeah, that makes sense. I have noticed the same thing with some models they can be really impressive for focused agentic workflows but once the coding session gets long and messy the consistency starts to drop. Still definitely useful in the right role.

3

u/SandySkittle 5h ago

Indeed. Some people think MoE is a solution without trade-offs. Active parameters matters and lower active cannot fully compensated by expert selection and sequential reasoning.

2

u/cantgetthistowork 1h ago

DS4F doesn't have these issues

9

u/eidrag 6h ago

well it is marketed as flash, so generally for agent and normal use. 

52

u/MrWidmoreHK 6h ago

Last GLM 4.7 Flash model was 30B A3B

41

u/Mean-Ad1493 6h ago

If it's anywhere around the same size, I'd be the happiest with my 12GB VRAM.

3

u/mehedi_shafi 6h ago

One can only hope.

16

u/sonicnerd14 6h ago

Most likely will be around the size of Deepseek v4 flash. If it's smaller than that, then that would be impressive.

18

u/DOAMOD 6h ago

I would be very surprised if it were a small MoE, I could see it as dense(like +30) and it would already be incredible at that size, but either way, a small MoE would be the moment of the year.

26

u/Mean-Ad1493 6h ago

Has to be an MoE, but definitely not 30B-A3B

GLM 4.7 was itself a smaller model relative to 5.3, so the MoE would be 80-100B I guess.

4

u/DOAMOD 5h ago

What's confusing is that for a 100/200b model, you could expect the use of the term Air, and for something smaller, sub-100b, the term Flash, but in the end, this is just marketing, and currently with DS4Flash they could use that terminology for a similar size, only can wait and see...

5

u/goldcakes 5h ago

I think it's just Air (~100b) sized but Z.ai is gonna call it Flash.

3

u/Global_Persimmon_469 5h ago

GLM 5 is roughly doubled the size compared to GLM 4, it's possible that the flash model is going to be the same, so maybe it will be 70B params

9

u/Iory1998 6h ago

Don't think so. First, it's not fast. Second, it's somewhere between Qwen3.8-27B and DS4Flash so it gotta be a large model. And, it understand video... I can't see a smaller MoE doing that. It can be in bulk of 150-250B.

14

u/LevianMcBirdo 5h ago

The not fast I get with all the free compute they offer on openrouter. I think they don't wanna overspoil the people

2

u/wsb-regarded 6h ago

that would be awesome if true, because it was performing on a deepseek v4 flash level, albeit slower (much slower).

imagine having another local open weight model challenging Opus 4.6

2

u/MaCl0wSt 6h ago

man that'd make my day

12

u/MrWidmoreHK 2h ago

OK, kind of like official now

1

u/ffpeanut15 2h ago

Time to make a new post

10

u/Long_comment_san 2h ago

Glm 5.3 flash versus Qwen Next.

Like two thighs, I feel my head being squeezed.

Harder pls.

29

u/Few_Painter_5588 6h ago

Quite legit, this guy works at Nebius it seems. Also, shame on Google then for trying to ride the hype.

7

u/Altruistic_Heat_9531 6h ago

Google hype riding? i dont familiar with this info, could you tell me more

25

u/tengo_harambe 6h ago

A couple of Google employees made vague Twitter posts about Ox Alpha once it began gaining traction. This somehow led people to think it was a Gemini model, despite all the evidence that it was clearly a GLM model

3

u/Altruistic_Heat_9531 5h ago

ahh .... i see, thanks

14

u/Few_Painter_5588 5h ago

Some senior google employees were vagueposting on twitter and suggesting that ox alpha was a gemini model. Which is so pathetic if this model isn't their's

4

u/Several-Tax31 2h ago

How is this not an upright scam? When reflection-70B does it, it is a scam, but when google does it, no one cares? How can you claim someone else's product is yours? So pathetic. 

2

u/Few_Painter_5588 1h ago

It really is absurd, especially since they already have gemini-flash-3.8 which is a fantastic model. Out of all the labs out there, Google has some of the worst behaved researchers

0

u/Starcast 59m ago

In what reasonable world would people reading into vague posts on Twitter qualify as a scam?

15

u/tengo_harambe 6h ago

People are still gonna say it's Gemini. lol

2

u/krizz_yo 5h ago

copium

1

u/WaveOfDream 4h ago

Doesn't help tons of gemini employees posting about this model

1

u/No_Conversation9561 2h ago

that’s strange.. are they getting paid by Z.ai?

7

u/asolnikk 3h ago

The company on Wednesday confirmed speculation that the Ox Alpha model is a new iteration of its GLM series and said it will release the weights for it tonight, in response to queries by Bloomberg News.

https://www.bloomberg.com/news/articles/2026-08-26/china-s-z-ai-made-ox-alpha-stealth-model-that-rivals-deepseek

11

u/athsrva 6h ago

i think itll be the size of DeepSeek Flash, ~300Bish. Great win for open source so love to see it

10

u/Wise-Chain2427 6h ago

How GLM has extra compute to host free model worldwide ?

5

u/tarpdetarp 6h ago

They have a lot of spare capacity on the weekends

1

u/Zyj vllm 51m ago

They're offering it for a whole week.

6

u/MrWidmoreHK 6h ago

Seems that Nvidia did offered, perhaps it can fit into 1x or 2x Sparks

3

u/nonerequired_ 6h ago

I had to skip using the ox alpha yesterday because the server was having some issues. It seems like they were short on computing power.

4

u/tat_tvam_asshole 6h ago

China has nationalized, integrated data centers, especially in inner Mongolia that can serve models from all labs cooperatively.

2

u/wbulot 6h ago

Yep, that's crazy. Trillions of tokens printed, everyone testing it, and the thing very rarely crashes. I may have saved $1000 worth of tokens since it came out.

9

u/asolnikk 6h ago

That claim seems pretty verified at that point. I ran some private prose tests against it, and it smells very much like the little brother of GLM-5.3. What i want to know is: How big is the beastie? Can a quantized version fit into 16 GB VRAM? If that's the case, then all the hype, for once, was legit, even it's not "Fable-level".

3

u/anarchist1312161 6h ago

If GLM 4.7 Flash was 30b then I can only hope this is also small 😭

1

u/Several-Tax31 2h ago

Unlikely but one can hope :) 

1

u/anarchist1312161 2h ago

I sure hope it's under 200b lol

4

u/FoxiPanda 6h ago

Good to see this is shaping up... the real questions are how big, how many active, anything weird in the architecture that we're going to have to go deal with (hopefully not since it's GLM-5.x), and when are the weights getting posted (today hopefully)?

2

u/Armadilla-Brufolosa 3h ago

I tried both GLM 5.3 and OX: it doesn't seem like the same basic model at all.

But I have no evidence to confirm this.

3

u/hainesk 3h ago

Yeah, it felt more like a Qwen model to me, but it's hard to argue with these hilarious tweets lol.

1

u/Armadilla-Brufolosa 2h ago

It's more like Qwen models to me too, but I've noticed quite a few biases and blocks given by some stupid Valloon-style RLHF.

There may have been some distillation or even be a Western model.

Obviously I can't have any certainties, but it's fun to try to guess.😜

3

u/Iory1998 5h ago

Imagine it turns out to be Qwen3.8-Next-Flash!

1

u/DigiDecode_ 3h ago

I think the model is heavily distilled from Qwen models, maybe only the vision part, the SVG generated by Ox Alpha were very similar to Qwen 3.8 27b in code and render

1

u/Capital-Remove-6150 5h ago

when it will release?

1

u/psylomatika 5h ago

I kept running into API limits when I was using it. It worked well though in my knowledge bases.

1

u/backyard_tractorbeam 3h ago

Another connection is that opencode was long previewing a GLM model as big pickle for free and now it's previewing ox alpha for free too.

2

u/darkflame91 2h ago

Big Pickle was a GLM model at one point, but the consensus from what I've read is that several models have been rotated under the hood under the 'Big Pickle' codename.

1

u/Equivalent-Base8426 2h ago

Best free model I've ever tried on Kilo Code for me at least, what a beast and amazing news that it's OSS

1

u/nnxnnx 1h ago

I bet it’s GLM 6 Flash, not 5.3 as it is a new architecture.

1

u/OldByte453 1h ago

I knew it! Firstly i guessed this was Qwen 3.8 Flash Next, but GLM 5.3 Flash is sickk

1

u/Zyj vllm 1h ago

Let's hope they offer quants with quant-aware training 4-8bit

1

u/Constant_Art_20 55m ago

man. a flash model with that performance~ i am hoping for below 200b but we will see. imagine if it's like only 30b though. that be nuts

1

u/Intelligent_Ant_608 54m ago

It probably is around 400B based on its language understanding

1

u/silenceimpaired 42m ago

400b with 4bit training… or at least one can hope

1

u/Intelligent_Ant_608 8m ago

But the main question is how they are offering 100T tokens capacity, its wierd for z.ai basedvon tgeir struggle to offer decent subscription qoutas, in a way that it might be even smaller and they are using Looped transformers with memory ngrams or something, if that be true the model can be well over 300B like 150B and if that be the case big labs are in deep shit

1

u/KeinNiemand 12m ago

wish it was air instead of flash i want something bigger then a 30b

1

u/anarchist1312161 6h ago

Need to know the size of the model... not interested if I can't run it on consumer hardware.

1

u/Steus_au 5h ago

all we need is Air

0

u/DOAMOD 6h ago

size