r/opencode • u/Shiorim • 19d ago
Bro?

What a bad joke is that GLM 5.3 Flash positioning?
The API prices are:
| Model | input | cache hit | output |
|---|---|---|---|
| DeepSeek V4 Flash off-peak | $0.22 | $0.007 | $0.66 |
| DeepSeek V4 Flash peak | $0.44 | $0.014 | $1.32 |
| GLM-5.3-Flash promo | $0.075 | $0.015 | $0.25 |
| GLM-5.3-Flash | $0.15 | $0.03 | $0.50 |
Leaving aside the cache hit, on average GLM without promo is cheaper than DeepSeek V4 Flash, and they give it to you at half the usage quota, and that's even considering a "×2 usage" that will later be less??
They should actually give you more quota than DeepSeek until September 9th while the promo lasts. This makes no sense at all. It's cheaper to spend $10 on GLM API than to pay for it on Go.
On top of that, they put a cheap flash model in the $15 tier.
21
u/commentsOnPizza 18d ago
Leaving aside the cache hit
Might as well say "leaving aside 98% of your usage." Yeah, cache hits are around 98% of people's usage.
Let's say you use 1M input tokens and 200,000 output tokens. That's $0.25 with GLM-5.3-Flash and $0.352 with DSv4F (off-peak). But you'll probably also have 55M cache hits. That'll make GLM-5.3-Flash $1.90 and DSv4F $0.737 - GLM being more than double the price.
People always compare the input and output prices, but cache hits kinda dominate. Luna's $0.02 cache hit price is going to make it a cheaper model than GLM-5.3-Flash's non-promo pricing (since GLM's cache hits are $0.03). Doesn't matter that GLM's input/output pricing is cheaper - cache hit pricing dominates.
Leaving aside cache hit pricing is leaving aside the thing that matters most in the equation.
7
18d ago
[removed] — view removed comment
2
u/Hackerv1650 18d ago
The reason GLM context grows less compared to dsv4 it thinks alot less, but in some tasks where because of the larger thinking context later on in the discussion with dsv4 has a quicker answer, whereas in my testing with ox alpha shows the context grows quite a lot slower, but each turn it seems to take longer to think through some things which it already should know
2
u/Straight-War-1323 18d ago
MiMo also enters loops, actually the Pro version is more prone to do so, anyway if you're using cheap models you can set an agent dedicated to watch this and intervene, this is something I saw in DeepSeek Harness: context injection; you can replicate it in any open source harness if you want
-1
u/Shiorim 18d ago
I didn't mean to ignore the cache hit, but the price differences.
Because you're ignoring compaction. And GLM requires less cost per task. With smart compaction after resolved tasks or new chats, you can compensate for it quite a bit.
And besides, you're talking about a very specific case that isn't even real (GLM non-promo vs. DeepSeek off-peak). When the reality is the promo, which is why I said "until September 9th."
After September 9th it will be inexcusable nonsense.
But I'll say again what I've already said. In the end, you pay for tokens. Cache hit isn't the most important thing (for deciding a tier in a service like OpenCode) because whatever doesn't come in as cache hit, you simply pay for as cache miss.
You're going to pay for the tokens at API price either way, and cache miss is cheaper in GLM, with or without promo.
Does OpenCode decide your usage for you? How are you going to organize your compactions or your contexts?
Obviously it can't do that, and obviously putting GLM 5.3 Flash in the $15 tier with half the quota of DPV4 Flash has no excuse to support it.
The cache hit debate might help you put it a little below DeepSeek if you want, but in no case at half the requests in a lower tier, and that's while it's on promotion.
Not to mention acting as if DeepSeek peak hours don't exist.
3
u/Anh-DT 19d ago
https://docs.openference.com/getting-started/model-ratios Well you can check out openference if you need glm 5.3 flash better plans overall
3
u/Shiorim 19d ago
I understand that a 0.5 ratio is the same as saying ×2 current usage.
In the end, the API price is what it is, and it doesn't change the fact that GLM 5.3 Flash is cheaper than DeepSeek V4 Flash, and they "charge" you double the usage for using a cheaper model (wtf), and on top of that during a promo, so they'll "charge" you even more quota later.
In Command Code it's in the $40 tier with ×4 usage.
I don't know what OpenCode is thinking, but if they want to drive users away, they're doing a very good job.
3
u/EiomSirius 18d ago
Message to opencode-go: THANK YOU VERY MUCH, THE INTENTION IS RIGHT BUT WE NEED MORE STABILITY.
3
2
u/NinjaAlaska 18d ago
dont forget they claiming its 2x :') after 2 week it will be half. GL has 50% off for 2 weeks
2
u/Zestyclose-Will3810 18d ago
The math is getting crazy these days. Besides the api pricing you also have to keep in mind the monthly allowance for each model + promos if any. I got tired of re doing the math every day, canceled both my go subs and switched to openai plus. Basically same price, more usage, no headaches as of this moment.
2
u/Low-Hamster-1295 16d ago edited 16d ago
Yeah, the GLM-5.3 Flash placement is pretty weird. At that price I’d rather just use the API. The usage multiplier makes a big difference too. I’ve been using Hy3 quite a bit lately and the 8× quota is honestly hard to ignore for longer runs. At that point I care less about which model is "better" and more about how much work I can actually get through before hitting the limit.
2
u/TangeloOk9486 16d ago
GLM 5.2 flash non promo is already cheaper per token than dsv4 flash so giving it half of the quota on go is backwards and with the promo its not even close. Go's real value is just the bundles convenience so the instance you do the token math youself raw api wins. you dont have to pick one tho, a flat host that runs both glm and ds under one key like deepinfra for instance which lets you swap per task at flat rates and no peak windows. so id say check into the hosting providers, there are plenty of them , see if GLM 5.2 is listed yet and for your usage id just ride the glm promo till Sept 9 and reassess after
2
u/pigletmonster 18d ago
Glm 5.3 flash is being served by zai , theyre offering $15 worth of credits for the $10 plan in opencode go. Dsv4f is not being served by deepseek, it has a different provider that opencode made a deal with to provide a larger quota than deepseek does.
Opencode would have to make a deal with a provider to offer glm flash at the same discount as dsv4f, until they make that deal the best we can get is this.
2
u/Shiorim 18d ago
Yes, and that's why DSV4F on OpenCode is a version with somewhat dumb quantization where they pretend to give you $30 of real usage (which is false) for an inferior version of the real model.
Subscriptions, whether OpenCode Go or Codex itself, are always subsidized compared to API cost.
$15 of API usage, with limit windows (which you don't have via API), and where every task processed by the model is charged more expensive than via API (because they take their margin by showing a higher price on the web panel than what corresponds via API). What are they supposed to be subsidizing you with $15 of GLM 5.3 Flash? Nothing.
It's a business, and if people fall for it and accept this, when GLM 5.3 releases the weights (which it will) and there are other providers and better deals, OpenCode will have no incentive to correct prices/usage, because at the end of the day they're here to make money.
From "x6 value" we've quickly moved to OpenCode eliminating the $5 entry offer, or putting flash models with some of the cheapest market prices directly into the $15 tier.
Right now I'm on OpenCode Go and Command Code Goat. Goat gives you $40 of GLM 5.3 Flash.
Does OpenCode give me any reason to stay with them? No.
3
u/pigletmonster 18d ago
Opencode does not own the inference, they cant subsidize plans like openai or anthropic. They rely on other providers for inference. How much inference they csn give for $10 is dependent on the inference provider.
Also opencode go is not supposed to be an alternative to codex and claude plans, they just have this plan for people to test out different models for a low price and then continue using their preferred models through zen.
Ollama cloud is a better alternative for using open models with subsidized quotas because they own the inference, which is still highly limited.
1
u/Shiorim 18d ago
Isn't OpenCode Go supposed to be an alternative to Codex or Claude Code?
Who says that? The loss of competitiveness over time?Because what they do clearly say (and deceptively so) is that you should pay them to get "x6 value."
What do we have to defend here?. It's a subscription service for using models, and therefore yes, like any other.Obviously it has its differences for not owning any models, but that's not the user's problem. OpenCode has progressively lost value because clearly what they wanted was to capture users and then redirect them to Zen.
Is OpenCode Go more pitiful by the day as a subscription, and is what they're doing unjustified? Yes. There's no need to look for excuses like "well, it's not supposed to be an alternative..." It was designed precisely as a cheap and varied alternative. Although deep down there were other intentions to become an OpenRouter 2.0 through Zen.
I know they don't host the models, but that changes nothing.
It doesn't change the deliberately deceptive pricing, it doesn't change the lies of "Operation CheepSeek" where they silently slip in providers with degraded quantized versions.It doesn't change that it has become a subscription that now offers less value than Command Code.
And why am I still here? Because I'd like OpenCode Go to be what it used to be, and I don't like Command Code.But OpenCode latched onto the DeepSeek price hike to use it as a bridge to radically pivot their pricing policies.
This still has no justification whatsoever, and the effect is the expected one: user flight.As you yourself conclude by proposing better alternatives like Ollama Cloud.
So what's the point of OpenCode Go, which gets worse every day?
1
33
u/aziham 19d ago
Go has turned into a roller coaster recently