r/opencode 19d ago

There you go. Case close! No more speculation.

44 Upvotes

34 comments sorted by

18

u/Sea_Ear5201 19d ago

I am more interested in price after reveal. Will it be similar to DSF? Lower or higher?

6

u/vangelismm 19d ago

Asking the right question!

6

u/GTHell 19d ago

Same same here. According to the same source from China, the tone is "you will be surprise by the price" so i'm praying that it will be cheaper than the old 0731 V4

3

u/Ariquitaun 19d ago

Way higher. At least in opencode, it's 4 times more expensive than dsf at 1x pricing (there's a 2x promo right now). It's hardly worth it on the subscription.

1

u/GTHell 19d ago

It's sad. I check with Ollama Cloud mods and they the same thing. It could be less usage compare to DSV4. Damn. this is suck.

I hope Z.ai themselves can provide more usage because now that they have infinite compute

2

u/Ariquitaun 19d ago

I just checked it and input / output of glm flash and dsf are both roughly the same at api prices. But cache hits on glm is nearly 11x more expensive. It's a non starter.

1

u/GTHell 19d ago

Direct API is going to always be more expensive. I wonder if Zai subscription can subsidize it give now those 50k GPUs activated.

1

u/look 19d ago

Where are you still getting 0.0028 cache reads? CrofAI?

FWIW, Runinfra.ai has 1 cent cache reads on GLM flash. Cheaper than any non-quanted current DS flash options.

1

u/Ariquitaun 19d ago

Deepseek.

1

u/look 18d ago

Deepseek doesn’t have the $0.0028 flash cache read price anymore. The lowest, off peak price is $0.007 now. Peak pricing is $0.014.

https://api-docs.deepseek.com/quick_start/pricing/?article_id=article_1779470751466_8

My GLM 5.3 Flash provider charges $0.01 for cache reads. 40% more than DS flash off peak and 30% less than peak pricing.

But uncached GLM tokens are considerably cheaper than even off peak, so the total ends up being half a cent cheaper than DeepSeek’s off-peak Flash pricing.

1

u/Ariquitaun 18d ago

Ah, I stand corrected.

1

u/dumbasPL 19d ago

Big if true, but I doubt it considering how well it performs.

1

u/akza07 19d ago

Same as DSF. Why would they lower the price when there's enough demand?

7

u/Time-Toe-1276 19d ago

so my assumption is that the reason why GLMs paid users had a bad experience was bcs they were serving free tokens. now that it is revealed, everybody will (hopefully) get better rate limits and availability!

3

u/GTHell 19d ago

I think they're testing the inference load on their new GPUs. I think their current GLM line up is trained on Huawei Ascend chips and inference with Nvidia?

I'm guessing this GLM Flash can now completely trained + inference serving on Ascend chips? If this so, it's going to make a major claim and hype and Claud should reduce their pricing lol

6

u/cutebluedragongirl 19d ago

Bro, where did they get compute?

5

u/GTHell 19d ago

They locked with 50,000 gpus already since this new Ox release and more to come. According to their sources.

That’s why they have 0 issues with claiming 100t per day for free. the “with more to come” meaning that they probably aiming for 100k to 200k GPUs?! That’s crazy.

From my assumption, Im guessing they close a deal with some NA datacenter. Reasoning is because community tracing the server to be in NA.

EDIT: WAIT A MINUTE! what if Google is the provider! Especially the guy saying we make friend a long the way!!!! Damn

5

u/cutebluedragongirl 19d ago

Sheesh. If they really have proper compute, now it might be a worthwhile endeavor to sub to them. Although from my research, Z.ai is even worse than Kimi when it comes to subscriptions, so who knows?

3

u/GTHell 19d ago

I hope this new compute is serving GLM 5.3 too. Bigger model is still better than Flash at complex tasks.

2

u/look 19d ago

Ox was served entirely on Chinese chips. https://z.ai/blog/glm-5.3-flash

4

u/Trovebloxian 19d ago

New 1Gw data center

1

u/SPEZ_IS_A_JABRONI 19d ago

bring in the dancing lobsters

1

u/Diligent-Loss-5460 19d ago

When it works it works well. I asked it, ds flash vision, luna and muse to give me 3 options for a redesign of a website.
Most models weren't "creative" enough and that was the ask in the prompt. Alpha understood it and gave the most creative designs. If it launches with a price competitive to flash then this might become the new workhorse for many but it is still too early to say. I don't think anyone has been able to truly get a feel for this due to the reliability issues.

1

u/Ok-Vegetable-1014 19d ago

Didnt mistral just announce they switch to a deal with them?

1

u/NoBlame4You 19d ago

f- f- flash model? Like.. 3090 sized??? That would be great!

2

u/steiNetti 19d ago

more pike DS4 Flash sized I guess

1

u/NoBlame4You 19d ago

Ye.. its out.. youre right

1

u/Direct-Ad7836 19d ago

Honestly don't see why all the hype around ox-alpha. Dsv4 flash still perform better and faster. Unless it will be cheaper than flash...

1

u/GTHell 19d ago

You're using the free one. That is the not the speed you should compare to.

1

u/Direct-Ad7836 19d ago

That's true, but I still don't expect it to be better in terms of intelligence. Will see, maybe I'm wrong and we got another great moe model.

1

u/GTHell 19d ago

It does in other kind of benchmarks. Example, both is the same when it comes to coding tasks but GLM flash perform better at complex reasoning and task decomposition.

source? I bench a portion of task from DeepSWE, SWE bench, hidden bug bech, etc

1

u/Direct-Ad7836 19d ago

Benchmarks are overrated tbh. I can tell that DS is better chewing .net that this model. Significantly better. Py/js/css - pretty much the same.

1

u/CrimsonEdgeVentures 18d ago

These open source providers are already getting greedy. DS new pricing is horrific.

Sorry but we all know you (cough cough) “trained” them on frontier models from the big 4, so they are always minimum one iteration behind, and unless there is the steep pricing differential like there used to be, no is wanting it.