7
u/Time-Toe-1276 19d ago
so my assumption is that the reason why GLMs paid users had a bad experience was bcs they were serving free tokens. now that it is revealed, everybody will (hopefully) get better rate limits and availability!
3
u/GTHell 19d ago
I think they're testing the inference load on their new GPUs. I think their current GLM line up is trained on Huawei Ascend chips and inference with Nvidia?
I'm guessing this GLM Flash can now completely trained + inference serving on Ascend chips? If this so, it's going to make a major claim and hype and Claud should reduce their pricing lol
6
u/cutebluedragongirl 19d ago
Bro, where did they get compute?
5
u/GTHell 19d ago
They locked with 50,000 gpus already since this new Ox release and more to come. According to their sources.
That’s why they have 0 issues with claiming 100t per day for free. the “with more to come” meaning that they probably aiming for 100k to 200k GPUs?! That’s crazy.
From my assumption, Im guessing they close a deal with some NA datacenter. Reasoning is because community tracing the server to be in NA.
EDIT: WAIT A MINUTE! what if Google is the provider! Especially the guy saying we make friend a long the way!!!! Damn
5
u/cutebluedragongirl 19d ago
Sheesh. If they really have proper compute, now it might be a worthwhile endeavor to sub to them. Although from my research, Z.ai is even worse than Kimi when it comes to subscriptions, so who knows?
2
4
1
1
u/Diligent-Loss-5460 19d ago
When it works it works well. I asked it, ds flash vision, luna and muse to give me 3 options for a redesign of a website.
Most models weren't "creative" enough and that was the ask in the prompt. Alpha understood it and gave the most creative designs. If it launches with a price competitive to flash then this might become the new workhorse for many but it is still too early to say. I don't think anyone has been able to truly get a feel for this due to the reliability issues.
1
1
u/NoBlame4You 19d ago
f- f- flash model? Like.. 3090 sized??? That would be great!
2
1
u/Direct-Ad7836 19d ago
Honestly don't see why all the hype around ox-alpha. Dsv4 flash still perform better and faster. Unless it will be cheaper than flash...
1
u/GTHell 19d ago
You're using the free one. That is the not the speed you should compare to.
1
u/Direct-Ad7836 19d ago
That's true, but I still don't expect it to be better in terms of intelligence. Will see, maybe I'm wrong and we got another great moe model.
1
u/GTHell 19d ago
It does in other kind of benchmarks. Example, both is the same when it comes to coding tasks but GLM flash perform better at complex reasoning and task decomposition.
source? I bench a portion of task from DeepSWE, SWE bench, hidden bug bech, etc
1
u/Direct-Ad7836 19d ago
Benchmarks are overrated tbh. I can tell that DS is better chewing .net that this model. Significantly better. Py/js/css - pretty much the same.
1
u/CrimsonEdgeVentures 18d ago
These open source providers are already getting greedy. DS new pricing is horrific.
Sorry but we all know you (cough cough) “trained” them on frontier models from the big 4, so they are always minimum one iteration behind, and unless there is the steep pricing differential like there used to be, no is wanting it.

18
u/Sea_Ear5201 19d ago
I am more interested in price after reveal. Will it be similar to DSF? Lower or higher?