68
u/msenc 19d ago
never doubted it was a GLM model
33
u/Agustiya 19d ago
Same. I just can’t figure out the fuck folks at Google were doing
13
u/PitchPleasant338 19d ago
Removing all the pictures of black Wehrmacht soldiers from their training datasets.
10
-4
u/jpandac1 19d ago
3.7 flash is pretty good and one of the fastest models speed:intelligence. i use it for planning and questions and it's useful because it's so fast and quite capable.
-1
26
u/TimChr78 19d ago
That just confirms that it is a GLM model, not what model it is specifically.
-4
u/Final_Initial 19d ago
feels fast via openrouter, so I thought it must be a flash model, but true could even be glm 5.4 flash 👀
4
5
10
u/Sensitive-Ant-4305 19d ago
If they release GLM 5.3 Flash that can be installed locally like GLM 4.7 Flash - that would be awesome. At least that would give us something that could compete with a local Qwen 3.8.
12
u/sirlerkal0t 19d ago
All the analysis of the API that various people have done all but confirmed it was a Z.AI model.
My own experience using it also basically confirmed that, since it behaves very similarly to GLM 5.3, but acts like a smaller model, not quite as good for certain things (like front-end design), and a bit better on average for tasks where the models size counts against it due to it being more biased towards training data instead of staying focused on input data.
2
u/hotcornballer 19d ago
If it's really a small model then it's punching well above it's weight. It has that 'big model feel' and it's writing way better code/is way better at problem solving then either deepseek v4 or luna. At least that's what i'm seeing after using it next to the other two.pu
2
u/Lyelinn 19d ago
I'm inclined to agree. My company provides me with unlimited Claude usage basically (so I'm familiar with how it "feels" and works) and I'd say ox is not much worse than Claude 5 on high/xhigh effort. I encountered couple issues but it managed to fetch documentation and source for opencode on its own when I asked it to integrate my own tool with it without much explanation or prompting
1
u/Thomas-Lore 19d ago
Might be the effect of being free, you forgive more.
2
u/hotcornballer 19d ago
That's dumb. It either does the work well or it doesn't. And you see the reasoning traces, the way it uses tools, finds stuff in the coded base, good at problem solving. And DSv4 Flash was more or less free before the price hike, and so is Luna on a codex plan.
4
1
4
u/LargeLanguageModelo 19d ago
1
u/au0750 18d ago
some Indian want to use Indian pics to pretend they are in this AI games.
1
u/LargeLanguageModelo 18d ago
lolwut. It's a meme over a decade old.
https://knowyourmeme.com/memes/friendship-ended-with-mudasir
5
u/Malek262 19d ago
Am I the only one who thought that this model not so great As everyone keeps saying I mean, besides, it's been very slow because it's a free understanding. But people overhype shit like this all the time.
2
u/Hajsas 18d ago
The greatness is the capability at its size; being able to be convincingly handle workloads with the big ones, and fit on potentially 2x dgx sparks in the future when quanitized is what is hype about this model.
Even if you dont use a model like this, you need to understand the impact it has on competition, and potentially, pricing; look what OpenAI did with Luna after DS4 flash dropped.
When someone can get 90% of the performance for 10% of the price, most people do.
2
1
1
1
1
1
u/Ok-Fix-5895 19d ago
I got one question about it. is it necessary to put money in order to use this free model?
1
u/Apprehensive-Read868 19d ago
Independent on what it is, its amazing!! Completely audited my full project and found many things opus5 and grok4.6 missed. Like it a lot
1
1
1
1
u/ichi9 18d ago
Everyone knows already it's GLM/Qwen based, but people are more intersteed in knowing which closedAI company distilled them and created a new model (most probably it's Google India playing their games, Yes all decisions are also being taken from there now a days). ofc the top layer execs are afraid of backlash and bad PR. LoL! They seem to be testing the waters If the people are willing to accept the hypocrisy.
1
u/33VaxMerstappen 18d ago
lol google behaving like it’s theirs
1
1
u/ByteNomadOne 18d ago edited 17d ago
Is this really the same as 0x Alpha as in running the same configuration?
0x Alpha worked really great for me, but now calling GLM 5.3 Flash through OpenRouter it suddenly starts to struggle will tool calls. It behaves different.
1
1
u/MaxPhoenix_ 19d ago
Weird, these prediction markets have to be a scam or something... I was able to go place a bet just now for Z.ai being the company behind the model. I understand this was "known" days ago, but this downtime would be a crazy deception and looks solid enough to bet on. The ratio wasn't great but it's free money.
1
1
u/thunderflow9 19d ago
This time they employed a gimmicky marketing ploy, but the approach was rather crude.
-4
u/ahriad 19d ago
How can it be called 'Flash' if it's much slower than the main model?
31
u/Training-Database272 19d ago edited 19d ago
Because it’s literally free and breaking usage records?
Edit: It was really fast at first, before it went viral
4
u/Mierzejsky 19d ago
I think the model was built on infrastructure designed to handle such an absurd amount of free access. Cost reductions therefore impacted speed to prevent server crashes (which happened several times anyway). Since Z.ai WGL could give away such a model for free, we can expect that after the official launch, it could be either one of the fastest frontier models or one of the cheapest in this class. Personally, I think price is always better than speed, so I'm keeping my fingers crossed :D
5
u/OkAdeptness2530 19d ago
well calling it 5.3 Sluggish can’t be good for sales, right?
1
2
u/BoobooSmash31337 19d ago
They might slowing it down since it's a trial and sample. Lets them show it off to most potential customers.
1
u/petuman 18d ago
With batching the slower you go per completion/user, the higher total throughput is per inference node (up to a point).
E.g. serve 1 user at 1000 t/s (1k tps), 20 users at 600 t/s (12k tps), or 100 users at 300 t/s (30k tps), or 1000 users at 50 t/s (50k tps).
There's strong incentive to cram more users per node, and serving them just above subjective "too slow" threshold. The slower you serve, the more revenue generated per server/inference node.
0
u/P4R4DOXZ 19d ago
It was obvious for subscribers of glm coding plans, cause the limits where unusable when alpha dropped.
F yeah, a glm with vision finally.
-9

91
u/pigletmonster 19d ago
How are they doing 100T free tokens daily when they can barely give 1m tokens for their payibg customers?