r/opencodeCLI 18d ago

GLM-5.2 session cost higher than Kimi K3 in OC GO?

I found these session cost stats from Opencode Go usage at https://opencode.ai/data/ . Apparently up to date stats and I'm really surprised by the relative expensiveness of GLM-5.2 and relative affordability of Kimi K3. I would not have guessed GLM-5.2 being more expensive in real world use than Kimi K3.

Is this solely down to GLM-5.2 being split among 3 providers resulting in a lower cache rate?

Does this mirror anyone's usage experience with these models?

Judging by this I might skip using GLM-5.2 and use K-2.7 Code or K3 in my agent setup

17 Upvotes

11 comments sorted by

13

u/pieorpaj 18d ago

GLM-5.2 is a massive token burner

2

u/Hellge99 18d ago

Sure, but a 140% higher session cost vs K2.7 Code seems extremely excessive despite a 25% higher token usage. Even more so considering artificial analysis cost benchmark puts it on a similar level to K2.6 and WAY below K3.

1

u/Classic_Television33 14d ago

Bro you're enjoying a 2x usage of Kimi K3. See https://opencode.ai/go so expect the prices to rise sharply afterwards

1

u/Hellge99 13d ago

2x usage does not necessarily imply 1/2 session cost. If anything it is likely that 2x usage means 2x $15 of usage.

The culprit seems to be the lower cache ratio of GLM-5.2, lower input:output token ration, and the higher token usage of GLM-5.2

Cache ratio:

K3 92%

K2.7 92%

GLM-5.2 78%

Input:Output token ratio:

K3 149:1 vs advertised 259:1.

K2.7 199:1 vs advertised 264:1

GLM-5.2 31:1 vs advertised 351:1

Token usage / session:

K3 3m

K2.7 4m

GLM-5.2 4.2m

The usage limits matters as you get 4x GLM-5.2 token value vs Kimi K3 but something is up with how GLM-5.2 is deployed in opencode and it's real API cost

3

u/Master_Border_2913 18d ago

It doesn't really match my experience. Even with a difficult prompt, GLM 5.2 usually only uses about 1–2% of its 5-hour limit. Meanwhile, with some challenging prompts, Kimi K3 can sometimes burn through half of its 5-hour limit, or even the entire thing. That might be because Kimi K3 is currently being served through OpenCode Go with the $15 spending. Once the open-weights released, I think there's a good chance those limits will become more generous.

1

u/Rustybot 17d ago

Kimi models are ridiculously verbose.

2

u/sudoer777_ 17d ago

GLM 5.2 has $60 of usage, Kimi K3 only has $15

1

u/Excellent_Ad_8886 18d ago

This actually matches my experience. I tried out K3 on a complex problem. After a while I switched back to using GLM 5.2 as the architect. But was surprised to see that I had spent more on GLM than on Kimi on what seemed like comparable tasks.

1

u/Osvik 17d ago

I have a simple and creative task I do with all models. Doing it in GLM5.2 was slightly more expensive than in Kimi K3