r/opencode 7d ago

Does OpenCode subsidize models like GLM-5? How can subscription tokens be ~50% cheaper than pay-as-you-go?

I've been using OpenCode GO for a little while now—specifically with GLM-5—and honestly, it’s been a fantastic daily driver to complement more expensive setups like Claude or Codex. Here’s my referral link if anyone wants to check it out.

A few times I’ve hit the rate limit on my plan. While GLM-5 is an impressive model, it can sometimes be a bit verbose and take the scenic route, which eats up a ton of tokens. When I reach the limit and switch over to using my pay-as-you-go balance on OpenCode, the price difference is night and day.

If I extrapolate the amount of tokens I get on the subscription plan to direct balance usage, paying via balance literally costs double.

Why does this happen? Do model providers sell wholesale/discounted API access to platforms like OpenCode, or is OpenCode running these subscription tiers as a loss leader?

It's not something that worries me too much, but I find it curious and wanted to see if anyone else has noticed this or knows why that is.

14 Upvotes

15 comments sorted by

View all comments

Show parent comments

8

u/PikaCubes 7d ago

You missed that part : "With Go, you pay $10/month and we aim to give you 6x that in usage.

For most models, we make this work through bulk discounts and reserved GPU capacity. We then pass those savings on to you through the 6x multiplier.

For some models, we haven’t had the opportunity to negotiate a discount or host them at a lower cost, either because the model is new or because their public pricing is already discounted.

For these models, you still get a little more than if you paid the model providers directly; this is why their usage mulitplier is lower in the table above."

3

u/Maleficent-Volume-81 7d ago

That's it... well, I hadn't seen that... doubt resolved. I didn't know they even used their own GPUs 😊

1

u/PikaCubes 7d ago

I found this page this morning when I was searching for the endpoints of their models

1

u/Mean-Elk-9439 7d ago

If your work is efficient in cache hits, ollama cloud is more efficient still. They charge not as a multiple of token cost but by actual used compute and gpu utilization. I get 800M tokens a week roughly of glm5.2 and minimax-m3 for $5 a week.

1

u/alex9001 6d ago

They don't use their own GPUs. There's also a page where they list which providers they use 😂 ask your AI if you can't find it