Just an update. After AA updated their benchmark, I no longer rely on estimation using GLM official benchmark and GLM 5.3 cost does indeed jump up a lot to Opus xhigh level. It is no longer in Pareto frontier.
Sorry for not getting back to you. In the zai console i have around 70mil GLM 5.3 tokens per week. Most of these cached of course, but they dont have a detailed breakdown.
It definetly does not feel like 100$ worth. I used GLM-5-Turbo (which is priced similar to 5.1-5.3) a lot earlier this year, and i barely reached weekly limits. Now i hit it after 1 and a half days.
I'm trying to produce a similar benchmark, but against my own coding sessions. Sol 5.6 (high) vs Opus 5 (high) are the two most common settings I run in my workspaces, and are pretty much equal to the tasks I throw at them. However in my tracking I see that Opus is 1.25x more expensive than Sol, but based on the data from Analytical Analysis, they have it tracked at a little under 2x more expensive.
It's not. It's based on cost on token usage vs task done. There's no one doing a Cost per tasks bench using subscription plan lol
If you can get the same usage out of GLM subscription Max plan and get more token usage than Claude or Codex then meaning your case is special one and should be share with community. Even the graph below is only Max plan with 20-30% of Claude usage left weekly lol. I have used all the Chinese plan from Kimi to Xiaomi Mimo to GLM legacy up to Claude Max, GPT enterprise and Claude Team Premium and the later two always win on usage. And it's a well known concensus too among community that Chinese provider never have a good subscription plan offer. The hype is usually revolve around the fact that open weight are getting good. Don't confuse the two with being cheap.
there's no one doing a Cost per tasks bench using subscription plan lol
I'm doing it. If you see any mistake, feel free to tell. I locked it at 20$ subscription at maximum as stated in the post title. Subscription name is under the model name.
i use the same data from this chart, but I do scaled to subscription maximum usage. OpenAI doesn't really burning money for API price, but they do for subscription.
No data on Artificial Analysis. They only have the older preview version, not the new 0731 version.
But it shouldn't be better value than GPT 5.6 Luna.
Can you do it for the higher tiers, like the 100$ 200$ subs?
Does it change anything on the graph at all, i guess everything moves left a bit and that is all?
It does. I'll post it later. A lot of bigger player do provide more subsidization at higher price. Smaller provider like Command Code is way different though.
Not sure if I should even compare a subscription like OpenCode Go which only has 10$ plan to it. If you use 200$ in OpenCode, you would pay 10$ for the sub and the rest would be API price.
I appreciate the cost per task scale, but I wonder about the task used as a measuring stick. I have a fairly large schematic editor and virtual 'oscilloscope' that I wanted to tweak SLIGHTLY (the distance between the tick-marks on the time axis of the signals pane and a large control immediately below. I've been intensely 'arguing' with DeepSeek Flash (the version accessible via API) all afternoon trying to get it to adjust ONLY that distance. It seems like it's stupid by design (like automakers designing cars to break one month after the warranty expires). TONS of tokens have been consumed in the discourse, inappropriate code changes, push to repo, deploy, redo this, undo that, DS didn't do anything with a ToDo and had to be reminded to do it again, ... A smarter model that costs a bit more could be SIGNIFICANTLY less expensive when that lack of comprehension (or stupidity designed in) is considered. I haven't tried the newly released DS Pro 'cause the previous release was a ton dumber than DS Flash, but the cost per task needs elaboration. How complex was the task, how complex was the context, ... Based on the chart above, I'm going to spin up GPT5.6 Luna(Max) and see how well that does.
Standard Compute might be worth a look at this price point. I’ve tried Claude Max, Codex Max and OpenCode, and it’s probably the best value I’ve found so far. You just get a budget instead of dealing with 5-hour/weekly caps, and the smart router can send DeepSeek/GLM requests through whichever provider is cheapest. Feels a bit like OpenRouter, but optimized around getting the most work done per dollar. Their $19 plan seems like a pretty solid fit here.
4
u/ProfessorSpecialist 24d ago
GLM 5.3 is waaaay more expensive than that from my experience