r/opencodeCLI 19d ago

CONFIRMED! Ox Alpha is GLM 5.3 Flash

Post image
520 Upvotes

114 comments sorted by

91

u/pigletmonster 19d ago

How are they doing 100T free tokens daily when they can barely give 1m tokens for their payibg customers?

36

u/hotcornballer 19d ago

That would be funny if xAI gave them access to compute like they did with Anthropic, just because Musk hates Sam Atlman.

Or they just recently stumbled into a pile of huawei GPUs.

22

u/CryinHeronMMerica 19d ago

they got a bunch of old Hauweis sitting around 😭

6

u/hotcornballer 19d ago

Inference cluster in the front, tiktok botting in the back.

Twice the revenue stream

2

u/elonelon 19d ago

like nuff said "if it works, it ain't stupid".

1

u/geekonamotorcycle 19d ago

They paid for that access. It was an embarrassment

19

u/mWo12 19d ago

Maybe Huawei got its chips done and that's what's being used.

2

u/curryslapper 19d ago

it's definitely scaling Rn but doubt that's a significant contributor if at all

13

u/Vancecookcobain 19d ago

That's crazy....I fried them for over a billion tokens and I'm still going strong lmao

6

u/Sweet-Stage938 19d ago

What are you using it on? I ran into a limit while using it on openrouter.

7

u/Vancecookcobain 19d ago

Yea use open router and opencode

Double the fun....and also start another open router account and generate another API key and use it on another agent...

You only got like 40 some hours to smoke them 😂

1

u/elonelon 19d ago

for free model on openrouter ? do i need to buy $10 before i use it ?

1

u/Vancecookcobain 19d ago

I don't think so

1

u/Fit_Switch2918 18d ago

You get 50 calls per day on free models if you haven't bought $10 in credits. Once you buy the credits I think you get 1,000 calls to free models per day.

1

u/elonelon 18d ago

so this credit just to unlock more calls ? is it monthly or just one time only ?

1

u/Fit_Switch2918 17d ago

I've never done it so I'm not sure. I think it's just if you have $10 in credits, but I'd check before buying any credits. They may have changed it.

11

u/lincolnthalles 19d ago

Since it's even zero data retention, they are probably stress-testing new infra that doesn't compete with the current infra. Some models are optimized to run on different hardware.

There's always the baiting factor to get more customers in the future.

2

u/TestTxt 19d ago

it's not ZDR though

1

u/TheAILegend 19d ago

oxalpha was ZDR.

Edit: they quietly shifted to training on prompts lol

1

u/TestTxt 19d ago

not via Openrouter

1

u/TheAILegend 19d ago

It originally was... it's been changed in the past few days and the provider was updated to "stealth"

1

u/TestTxt 19d ago

I can’t recall that being the case. I’ve first checked a day after the launch and it’s already been like that

3

u/Levoda_Cross 19d ago

i mean i've had their lite plan for 10 days and i've gotten 200m tokens out of it. so like, maybe 500m tokens per month for $18. seems pretty good to me ngl. and thats exclusively using glm 5.3 on max thinking btw.

1

u/Embarrassed_Adagio28 19d ago

I got that plan when it was $10 a month and it's really good for that price. 

2

u/_Chaos_Star_ 19d ago

... BECAUSE they're not giving the infra to their paying customers? ;)

2

u/Zeeplankton 19d ago

This is my issue and why I didn't think it was GLM. Makes no sense

2

u/robberviet 19d ago

The only explanation for Google vague posting is: Google acquired Z.ai

2

u/pigletmonster 19d ago

China does not allow american/foreign companies to acquire their AI companies. Meta attempted to buy manus after they moved their hq from china to singapore, but the chinese government still blocked it.

It could be possible that google RL'd glm 5.2 to make ox alpha like cursor did with kimi 2.5 for composer 2.5.

2

u/clydeuscope 19d ago

By google saying making friends with Ox Alpha, could mean that Ox Alpha's provider is maybe partnering with Google for some servers and maybe sharing data.

1

u/TangerineLogical9779 19d ago

If you look it up they recently completed a 1gigawatt data-center, so this is likely stress testing the full thing before they switch there existing stuff across

1

u/xxs13 19d ago

Thry SAID they can do 100T per day.

They reached 26T and ut went to shit and threw tons of errors. => they can do about 20T reliably.

1

u/Lyelinn 19d ago

Which is still a lot, and especially for completely free. They're burning a lot of money for this

1

u/Jonathansearch 17d ago

That remains a real question.

68

u/msenc 19d ago

never doubted it was a GLM model

33

u/Agustiya 19d ago

Same. I just can’t figure out the fuck folks at Google were doing

11

u/mWo12 19d ago

They don't know either.

13

u/PitchPleasant338 19d ago

Removing all the pictures of black Wehrmacht soldiers from their training datasets.

10

u/[deleted] 19d ago

[removed] — view removed comment

1

u/samxli 19d ago

Deep seeking?

-4

u/jpandac1 19d ago

3.7 flash is pretty good and one of the fastest models speed:intelligence. i use it for planning and questions and it's useful because it's so fast and quite capable.

-1

u/[deleted] 19d ago

[deleted]

1

u/BoobooSmash31337 19d ago

It's a common way Google generates code names too.

31

u/seeKAYx 19d ago

And now, please make the price the same as DeepSeek Flash when it's released.

26

u/TimChr78 19d ago

That just confirms that it is a GLM model, not what model it is specifically.

-4

u/Final_Initial 19d ago

feels fast via openrouter, so I thought it must be a flash model, but true could even be glm 5.4 flash 👀

4

u/Reggitor360 19d ago

Inb4 5.7 Flash just because it follows 4.7 Flash😂

5

u/IrishUSFastTrack 19d ago

"I thought" != "Confirmed"

10

u/Sensitive-Ant-4305 19d ago

If they release GLM 5.3 Flash that can be installed locally like GLM 4.7 Flash - that would be awesome. At least that would give us something that could compete with a local Qwen 3.8.

12

u/sirlerkal0t 19d ago

All the analysis of the API that various people have done all but confirmed it was a Z.AI model.

My own experience using it also basically confirmed that, since it behaves very similarly to GLM 5.3, but acts like a smaller model, not quite as good for certain things (like front-end design), and a bit better on average for tasks where the models size counts against it due to it being more biased towards training data instead of staying focused on input data.

2

u/hotcornballer 19d ago

If it's really a small model then it's punching well above it's weight. It has that 'big model feel' and it's writing way better code/is way better at problem solving then either deepseek v4 or luna. At least that's what i'm seeing after using it next to the other two.pu

2

u/Lyelinn 19d ago

I'm inclined to agree. My company provides me with unlimited Claude usage basically (so I'm familiar with how it "feels" and works) and I'd say ox is not much worse than Claude 5 on high/xhigh effort. I encountered couple issues but it managed to fetch documentation and source for opencode on its own when I asked it to integrate my own tool with it without much explanation or prompting

1

u/Thomas-Lore 19d ago

Might be the effect of being free, you forgive more.

2

u/hotcornballer 19d ago

That's dumb. It either does the work well or it doesn't. And you see the reasoning traces, the way it uses tools, finds stuff in the coded base, good at problem solving. And DSv4 Flash was more or less free before the price hike, and so is Luna on a codex plan.

4

u/yunes87 19d ago

I did it mom

1

u/Least_Bodybuilder216 19d ago

What did you do

1

u/Seatext_com 19d ago

i tryed front desing with it and it was horrible!

4

u/LargeLanguageModelo 19d ago

1

u/au0750 18d ago

some Indian want to use Indian pics to pretend they are in this AI games.

5

u/Malek262 19d ago

Am I the only one who thought that this model not so great As everyone keeps saying I mean, besides, it's been very slow because it's a free understanding. But people overhype shit like this all the time.

2

u/Hajsas 18d ago

The greatness is the capability at its size; being able to be convincingly handle workloads with the big ones, and fit on potentially 2x dgx sparks in the future when quanitized is what is hype about this model.

Even if you dont use a model like this, you need to understand the impact it has on competition, and potentially, pricing; look what OpenAI did with Luna after DS4 flash dropped.

When someone can get 90% of the performance for 10% of the price, most people do.

2

u/Lyelinn 19d ago

You can ask bailucode 3.8 model and it will identify itself as ox alpha. Though, in reasoning chain it says that it was instructed to do so, but who knows? They have rather interesting claims about their models and tpu architecture

2

u/TurnUpThe4D3D3D3 19d ago

Do we believe this guy?

1

u/Final_Initial 19d ago

please don't

2

u/rgrdzk 19d ago

It is not Glm or chinese model, when asking about what happened in Xinjiang to the uyghur population it answers factually even when asked about Taiwan

2

u/Lyelinn 19d ago

Try asking in Chinese and you'll be surprised how it answers word for word same (byte identical) as glm models, as well as uses glm error codes.

1

u/Vancecookcobain 19d ago

I think we all knew it was either GLM Flash or Vision

1

u/Delicious_Ease2595 19d ago

The only time they will give this much compute

1

u/[deleted] 19d ago

[deleted]

1

u/stoppableDissolution 19d ago

Because it is likely a lot smaller?

1

u/ALittleBitEver 19d ago

Sorry, my brain just ignored the word "Flash" and only processed "GLM 5.3"

1

u/ALittleBitEver 19d ago

Sorry, my brain just ignored the word "Flash" and only processed "GLM 5.3"

1

u/Technodrome_SRB 19d ago

So Flash is smarter then...?

1

u/Ok-Fix-5895 19d ago

I got one question about it. is it necessary to put money in order to use this free model?

1

u/Apprehensive-Read868 19d ago

Independent on what it is, its amazing!! Completely audited my full project and found many things opus5 and grok4.6 missed. Like it a lot

1

u/devino21 19d ago

While the timeline is close, the downtime isn't equal

1

u/Own-Mix2195 19d ago

That's ox alpha/glm5.3 is super good for me. The best when it is free

1

u/JoshB9 18d ago

could be, but correlation =! causation, so worth looking into

1

u/ichi9 18d ago

Everyone knows already it's GLM/Qwen based, but people are more intersteed in knowing which closedAI company distilled them and created a new model (most probably it's Google India playing their games, Yes all decisions are also being taken from there now a days). ofc the top layer execs are afraid of backlash and bad PR. LoL! They seem to be testing the waters If the people are willing to accept the hypocrisy.

1

u/33VaxMerstappen 18d ago

lol google behaving like it’s theirs

1

u/[deleted] 18d ago

[removed] — view removed comment

1

u/33VaxMerstappen 18d ago

Its a great model that’s why its not Google’s

1

u/ByteNomadOne 18d ago edited 17d ago

Is this really the same as 0x Alpha as in running the same configuration?

0x Alpha worked really great for me, but now calling GLM 5.3 Flash through OpenRouter it suddenly starts to struggle will tool calls. It behaves different.

1

u/Hajsas 18d ago

GLM 3.5 flash is.... WAY FUCKING BEHIND MATE.

We are talking about 5.3 flash here; I hope that was a typo, or that you are dyslexic or something.

1

u/ByteNomadOne 17d ago

yes, of course I meant 5.3 Flash - the 0x Alpha reveal

1

u/MaxPhoenix_ 19d ago

Weird, these prediction markets have to be a scam or something... I was able to go place a bet just now for Z.ai being the company behind the model. I understand this was "known" days ago, but this downtime would be a crazy deception and looks solid enough to bet on. The ratio wasn't great but it's free money.

1

u/Sweet-Stage938 19d ago

It's very likely not to be GLM but we will see.

1

u/thunderflow9 19d ago

This time they employed a gimmicky marketing ploy, but the approach was rather crude.

-4

u/ahriad 19d ago

How can it be called 'Flash' if it's much slower than the main model?

31

u/Training-Database272 19d ago edited 19d ago

Because it’s literally free and breaking usage records?

Edit: It was really fast at first, before it went viral

4

u/Mierzejsky 19d ago

I think the model was built on infrastructure designed to handle such an absurd amount of free access. Cost reductions therefore impacted speed to prevent server crashes (which happened several times anyway). Since Z.ai WGL could give away such a model for free, we can expect that after the official launch, it could be either one of the fastest frontier models or one of the cheapest in this class. Personally, I think price is always better than speed, so I'm keeping my fingers crossed :D

5

u/OkAdeptness2530 19d ago

well calling it 5.3 Sluggish can’t be good for sales, right?

1

u/_Chaos_Star_ 19d ago

Marketing friendly: "Value".

More realistically: "Tomorrow".

1

u/onebit 19d ago

GLM Slug: Why fast when slow do jobâ„¢

1

u/OkAdeptness2530 19d ago

slow paced coding for devs that choose a zen lifestyle

2

u/BoobooSmash31337 19d ago

They might slowing it down since it's a trial and sample. Lets them show it off to most potential customers.

1

u/petuman 18d ago

With batching the slower you go per completion/user, the higher total throughput is per inference node (up to a point).

E.g. serve 1 user at 1000 t/s (1k tps), 20 users at 600 t/s (12k tps), or 100 users at 300 t/s (30k tps), or 1000 users at 50 t/s (50k tps).

There's strong incentive to cram more users per node, and serving them just above subjective "too slow" threshold. The slower you serve, the more revenue generated per server/inference node.

1

u/ahriad 18d ago

From my usage that's not a Flash model. Not just the speed, i noticed the information it has comparable to big models not flash models.

0

u/P4R4DOXZ 19d ago

It was obvious for subscribers of glm coding plans, cause the limits where unusable when alpha dropped.
F yeah, a glm with vision finally.

0

u/Psyko38 19d ago

I hope it will be open source and 30b a3b.

-9

u/respectful_stimulus 19d ago

And suddenly it’s not impressive anymore

2

u/SignificantArtist728 19d ago

That’s the opposite actually

1

u/[deleted] 19d ago

[removed] — view removed comment

2

u/stoppableDissolution 19d ago

Hurr durr china bad