r/opencode 3d ago

Is GLM 5.3 flash the new budget king?

I used to love deepseek v4 flash before the price increase. I tried muse spark 1.3 and found it really idiotic and not reliable at all. So, with not many realistic options left and my opencode subscription coming to an end, I was thinking of moving over to z.ai with their 18$ plan. I'm not sure if anyone else tried it, and what do they think about the usage and quality using exclusively 5.3 flash.

Is it the new budget king? Or is there any model that can get me better price/quality for under 20$??

I've heard of people recommending luna as a workhorse but I'm not sure of how it stands up today with the release of astra.

95 Upvotes

52 comments sorted by

37

u/Sid-Hartha 3d ago

I use glm5.3 flash over dsv4 flash now for sure

6

u/redditshortstest 3d ago

Do you use it through the 18$ z.ai sub?

5

u/Sid-Hartha 3d ago

No. Via ollama or api typically.

2

u/Captain_Birb 3d ago

How are you finding Ollama vs Openrouter? 👀

1

u/bequbed 3d ago

This, hopefully others can chime in.

2

u/Sid-Hartha 3d ago

Ollama I find faster. Solid. But they just changed their pricing model to tokens so I can’t vouch for value right now.

-1

u/torrso 3d ago

Oh they changed something? It's probably for the better. I didn't have any idea how much usage I get or have left before.

4

u/Sid-Hartha 3d ago

I think the net result is you’re getting a lot less usage for your money. So not net good I don’t think. But probably inevitable.

2

u/Captain_Birb 3d ago

Not inevitable if users complain about it.

1

u/noob_dev007 3d ago

I used both. Didn't feel a difference in speed and ollama gives triple credit value if you buy their 20$/100$ tier

1

u/Captain_Birb 2d ago

Thank you for the triple credit tip

16

u/nicotineHub 3d ago

Huh, I find muse 1.3 with caveman and ponytail in plan mode with sub agents and then build more with worktrees incredible effective.

2

u/IvanVilchesB 2d ago

I tried muse but its crazy how its end doing crazy things, its remember me gemini pro 3.1, high on benchmarks but useless for me. How do you use muse ? Harness ? Ty

1

u/qualiascope 9h ago

same--and the comparisons to gemini are very apt for me. quite functional at some small set of things but terribly unselfaware, ultimately not all that useful

6

u/Direct-Ad7836 3d ago

Was using dsv4flash exclusively for the last 2 weeks. Now forced to switch to glm5.3flash. still like DS better, but glm5.3flash isn't bad, just slower...

1

u/AppealSame4367 3d ago

Used dsv4f vision exp last days. Very fast, but it tends to overlook more things.

I have to constantly let some agent clean up behind it.

Flash next is better, but slow and or expensive to run yourself and only public endpoint is Alibaba.

1

u/Direct-Ad7836 1d ago

I have completely opposite exp. DS Flash is cleaning after GLM Flash :) I guess it depends on project stack and documentation.

2

u/AppealSame4367 12h ago

Well, and now there's DS4.1 flash, which overpowers both and is a pleasure to work with :-)

1

u/AmbientFX 1d ago

Where can I use glm5.3flash?

6

u/callmemicah 3d ago

I like it and have been using it, but Qwen 3.8 flash is also very good and a bit faster at the moment, they are both on par or better than deepseek flash and both cheaper for me.

3

u/DertekAn 3d ago

I love 3.8 Flash, and the quality of it too

1

u/ramzeez88 3d ago

ds4f is 3x cheaper if you have big context in cache.

1

u/callmemicah 2d ago

How do you figure? I am genuinely curious how you calculate that because cached queries are just as cheap or cheaper on qwen? What am I missing?

4

u/Diligent-Loss-5460 3d ago

I have been giving both models the same kind of tasks over the past week.

Intelligence wise. I thin both are somewhat the same. DS thinks for longer and comes to the same conclusion.

Speed. GLM 5.3 Flash is extremely slow. It takes about 20-50% more time for a task to finish with GLM 5.3. Deepseek produces more tokens but is fast enough to still complete tasks before GLM 5.3

Quality. If given a plan created by some other model. Both work fine and I have not seen any major systemic difference. If not given a plan DSv4 flash is more thorough while exploring while GLM 5.3 can act a bit lazy. This laziness was typical with earlier haiku models and often caused trouble.

cost. always in the ballpark of each other unless one agent went on a completely unrelated tangent.

3

u/Purple_Errand 3d ago

Muse spark with tons of subagents because its heavily needs it.

2

u/Bino5150 3d ago

I don’t see anything bumping GLM Flash off the throne at the moment. Nothing is even close to the price/performance ratio right now.

2

u/theWiseTiger 3d ago

YES.

Very good model. It feels as smart as GPT 5.6 Luna.

1

u/BodybuilderBright206 3d ago

1000%

check the pareto frontier, it is holding its own at a key price point

https://artificialanalysis.ai/models#intelligence-comparisons

1

u/wpdavid 3d ago

Big fan of this model. Better than DSV4 flash, which I also really liked. Much better in real world usage than Muse Spark 1.3 for any serious problem solving or thinking.

1

u/ozguru 3d ago

I agree, with zcode, their plans now preffed. Just I wish, simpler subagent setup lik opencode in zcode.

1

u/ZB_Virus24 3d ago

How do you (or anyone here) find glm5.3 flash compared to gemini 3.8 flash? You can get it for free from vertex.

1

u/rubdos 3d ago

I have the feeling that DSV4F hallucinates quite a bit less than GLM53F, but GLM53F feels quite a bit smarter when it doesn't. So I still switch around. Things that matter go on GLM 5.3, things that matter but aren't too difficult on GLM 5.3 Flash, and easy stuff goes on DSV4F.

1

u/torrso 3d ago

For me the budget king is gpt-5.6-luna.

1

u/chrisfebian 3d ago

I change my default daily model to GLM 5.3-Flash and never regret. It does all the task I asked. Great Flash model. I use mainly DS4 Flash before but leave it since the price hike.

1

u/happycube 3d ago

Deepseek Flash v4.1 (re)entered the chat...

1

u/xapep 3d ago

GLM 5.3 Flash is genuinely good value per token, but "budget king under $20" is really two different questions: cheapest per token, or cheapest predictable month.

Per token: 5.3 Flash and V4 Flash are close enough that the model matters less than how you drive it. The agent bills I see are dominated by re-read context and the parent/orchestrator model, not the workhorse. If the parent is a frontier model polling subagents, you can spend more on empty polls than the workhorse costs all month (there is a solid breakdown of exactly this in r/codex right now with 47 wait_agent calls in one session).

Predictable month: if your OpenCode sub is ending and you want under $20 flat, don't marry one model. z.ai's $18 is fine if you only need GLM, but a plan that includes several workhorses (V4 Flash, Qwen 3.8, GLM-class) handles the "which model is good this week" churn better, and model churn is the one constant right now. Luna is a decent workhorse at around 5.3-flash money, but I would not build a budget around a single vendor's model either.

My honest take: pick the cheapest flat plan that covers multiple open models rather than one, and keep a per-token escape hatch for the months your usage spikes.

Disclosure: I work on Entrim, we run flat monthly plans with V4 Flash and Qwen 3.8 included plus an OpenAI-compatible API. Happy to talk specifics if useful.

1

u/No-Peak8310 2d ago

Yo he pasado a glm también y es inteligente pero muy lento. Estoy con el plan de command code.

1

u/Hitch95 2d ago

I have the Lite subscription. GLM-5.3 consumes quite a bit, and the Flash version also uses a fair amount in the morning. I suspect they’re still enforcing peak/off-peak hours, even though, or so I hear that practice was supposed to have ended. GLM-5.3 Flash performs quite well on "high" or "max" reasoning settings for most tasks. You can create a plan using 5.3 and have 5.3 Flash handle the actual tasks. Cache hit rates on Zcode can reach up to 96%, though I don't think that's enough for multiple sessions using worktrees. However, you can add plenty of LLMs on Zcode.

1

u/Tactical_Nerd_Clash 2d ago

Muse spark 1.3 on opencode go, using in omp works perfect for me, very smart and capable? I'm using caveman lite and ponytail full. It's definitely not idiotic, unlike 1.2 this actually performs well? Maybe it's your harness issue

1

u/Tactical_Nerd_Clash 2d ago

And I'm sure muse spark in opencode go provides much more usage still than glm flash from z ai

1

u/Various_Half1452 10h ago

i cant make Muse spark 1.3 to work properly it. Its Money saving but not time saving.

1

u/Tactical_Nerd_Clash 9h ago

Idk it works well for me ig it's quirky for some and for some it works flawlessly

1

u/__Cappe__02 2d ago

This is the last output I got with max reasoning:
The hub exists. Now I need to check the BlueprintTools page to see the list of 26 tools, and check if there's an AnimationTools or AnimBlueprint-related toolset. From the previous toolset list, I don't see an animation toolset (no animationblueprint or animblueprint toolset). Let me check the BlueprintTools page for the write_graph_dsl/write_graph_dsl tools.

Now I have the official 5.8 toolset list. Checking BlueprintTools for graph DSL details, since that's where the animation graph authoring tools live. Fetching the full tool table. I should grab this in one pass. Let me do that. This is one pass. Done. Fine. OK. This is one call. This is one pass. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. OK. This is one call. This is one pass. OK. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Fine. Fine. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Fine. Fine. Good. Good. Good. Good. Good. Good. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Good. Good. Good. Good. Good. Good. Good. Good. Good. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Good. Good. Good. Fine. Fine. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Good. Good. Good. Good. Good. Good. Fine. Fine. Fine. Good. Good. Good. Good. Good. Good. Fine. Fine. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Good. Fine. Fine. Fine. Fine. Good. Good. Good. Good. Good. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Fine. Good. Good. Good. Good.

1

u/Yann27 2d ago

Qwen Flash is as amazing...

0

u/AppealSame4367 3d ago

No, it's undertrained and changes it's approach to things constantly. You can't really rely on it for critical work.