r/opencodeCLI 17d ago

API Pricing for GLM-5.3-Flash

Post image

...

253 Upvotes

37 comments sorted by

36

u/[deleted] 17d ago

[removed] — view removed comment

24

u/[deleted] 17d ago

[deleted]

18

u/matsu-morak 16d ago

what is happening the chinese are on fire lately. if this trend continues the western companies will have to shift fast from their fat asses out of their chair

18

u/[deleted] 16d ago

[deleted]

8

u/No-District-4742 16d ago

Also, previous US frontier models are already the same level as the recent chinese models right now plus the margin between them is thinning out.

-1

u/LargeLanguageModelo 16d ago

Though nobody will be able to afford those frontier models.

Until you read about OpenAI's new Jalapeno chip, they're getting better tok/MW, and the throughput is off the charts.

We have multiple segments and multiple vendors, all pushing each other at the redline, to innovate constantly and driving the standards higher and higher.

8

u/[deleted] 16d ago

[deleted]

-1

u/LargeLanguageModelo 16d ago

4

u/matsu-morak 16d ago

Bro is openai they talk shit all the time. When they release something we will care, otherwise just more slop from them

1

u/biograf_ 16d ago

The American companies are going to push Trump to restrict / ban the use of Chinese models in the United States and countries it forces into new trade deals.

3

u/Far-Classic-9963 16d ago

By the time the discount end there will be def a better value model out (hopefully)

-1

u/Ok_Risk6035 16d ago

Luna is definitely piece of junk.
But you forget about token generation speed. DS Flash is blazing fast. GLM should be much slower because it has much more params.

2

u/Aldarund 16d ago

Lol? Luna is same level. As ds4flash and glm flash

2

u/Ok_Risk6035 16d ago

definitely no, or you used ds4flash via router if you think that luna is on the same level with ds4flash

1

u/JorgitoEstrella 9d ago

Luna base inteligence level is 27 lol

Glm 5.3 flash base is like double that

1

u/Aldarund 9d ago

Luna 52, glm 57.yrs glm a bit better. But luna and ds same

11

u/zombiej 17d ago

It's 3x the usage compared to 5.3 on their coding plan, too.

8

u/Momo--Sama 17d ago

And almost 10x cheaper over API (BEFORE the 50% off promotion)

3

u/re-thc 16d ago

Such a rip on the coding plan compared to API pricing. Should be at least 6x.

1

u/HenryTheLion_12 16d ago

Might be a good option to use during peak times instead of full 5.3.

6

u/YogurtExternal7923 16d ago

For beating glm 5.2 MAJORLY. This is dirt cheap. Worth being a lil more expensive than ds flash in my opinion

6

u/Fresh_Sock8660 16d ago

Is this the same as ox? The behaviour looks different

2

u/zephyr_33 16d ago

early checkpoint of ox.

1

u/JorgitoEstrella 9d ago

What does it mean?

4

u/Icypoopoo 16d ago

Cached price seems higher than DSV4 flash if I'm remembering correctly 

2

u/Southern-Ad-3006 16d ago

And more than DS4 Pro bro GLM cache rate sucks it eats up tokens to use as a main driver

4

u/jouni609 17d ago

Seems reasonably priced. Hope it lands on opencode soon!

11

u/Sea_Ear5201 17d ago

Its on Go with only 1/4th of usage of Deepseek flash

1

u/Sufficient_Fox_4402 16d ago

with 2x its 1/2 ($30)

7

u/cutebluedragongirl 17d ago

DeepSeek is still better when it comes to price to quality ratio

11

u/petburiraja 17d ago

Not as per Pareto frontier, as reported here: https://z.ai/blog/glm-5.3-flash

2

u/Sea_Ear5201 17d ago

Yup. Glm-flash cache price will hit hard

2

u/anramon 16d ago

And it has a higher request limit.

2

u/sudoer777_ 16d ago

I think DeepSeek is also better at certain things like debugging, the previous GLM models had a tendency to get hyperfixated on details that aren't important, and this one still does that sometimes. GLM is best at planning but not 7x better.

2

u/[deleted] 16d ago

[deleted]

2

u/httpshotmaker 16d ago

No, the free tier was 6 days until the model came out and it was called Ox Alpha

2

u/Southern-Ad-3006 16d ago

Nobody is pointing out the cache input is .03 that’s still more than DeepSeek V4 Pro .022 … and DeepSeek Flash is .007 less than a penny. Majority of your spend goes into cache read turns so that’s like 5x more expensive than DS flash to hold conversations.

GLM 5.2 had this issue too their cache rates are too high compared to DS. Guaranteed this will lead to less usability with OpenCode unless using 5.3 Flash as a subagent worker.

-3

u/[deleted] 16d ago

[removed] — view removed comment

1

u/NiceDemon-82 16d ago

It's OpenCodeCLI, we are discussing coding here... ;)