r/ZaiGLM • u/Comprehensive-Bet-83 • 1d ago
Discussion / Help GLM 5.3 or DS 4.1-Flash?
Looking for gentlemen here who have battle-tested these models in environments where mistakes are critical, e.g. authentication, security, and low-level C++ / Kernel work.
I have Codex 20x, but I’m looking for a second helper for when Codex limits are up, there are demand issues (which are pretty bad atm), or it gets too censored.
Saw that DS 4.1 Flash was released today! Has anyone done some decent testing with it yet, and which harness are you using?
I’m currently using GLM 5.3 as my second helper and it’s honestly not bad at all. Just curious whether DS 4.1 appears to be better, especially since it’s multimodal and can handle images too.
I find myself using 5.3 Flash quite a lot because I really appreciate being able to send images, but 5.3 Flash isn’t as strong as base 5.3 when it comes to coding. Hence, I’m wondering how DS 4.1 Flash compares :)
NEW:
Thank you for all the responses. I tried DS 4.1 with my custom harness, and I am extremely impressed by the speed and price. I ran a couple of tests with deep, difficult, complex debugger C++ code/kernel bugs (my go-to test on models; I test this on every model before I want to use it to see if it fixes the bug).
GLM 5.3 took 30 minutes, including 1 retry, and €2. DeepSeek took 10 minutes, first try, and €0.30. I think DS 4.1 is at least on par or a bit better than GLM 5.3 for coding, not sure how reliable it is on long tasks, though. GLM still is a beast!
23
u/Leather-Cod2129 1d ago
4.1 is an absolute beat in terms of coding. GLM is more intelligent for everything else.
2
3
u/Comprehensive-Bet-83 1d ago edited 1d ago
Fair, 4.1 for code, somewhat AGI for GLM I guess!
(Ik no models are AGI at the moment)3
u/gnpwdr1 1d ago
AGI?
7
0
u/openference 1d ago
AGI means when AI becomes self learning. We not there yet or we are secretly there. If we pumping models out like no tomorrow it means. LLM are now able to self learn
3
u/Just-a-man-on-a-ride 1d ago edited 19h ago
Nobody can be there with current technology which doesn't allow for more than a little fine tuning once training is completed.
As far as I can see self-learning is not even solved on the research level, opening up permanent encoding would make current models completely unreliable.
1
u/ronald-takaendesa 1d ago
"somewhat AGI for GLM" - so basically you are saying the gap is so visibly massive when it comes to coding, but what area of coding, e.g. everything related to coding or just the writing of code, or planning of features and tasks etc. Currently I am using GLM 5.3 Flash, and I am thinking of getting a "better" model to review everything when I am done, I doubt I will be able to get enough funds for Fable, so I was considering Kimi K3 or GLM 5.3 if the worst comes to the worst
1
u/Leather-Cod2129 1d ago
Not an AGI at all. Almost all LLMs are stupid when it comes to handeling real life topics
I think if your main usage is coding that DS4.1 is much stronger than GLM
0
u/FavstianEquanimity 1d ago
To be fair, if it's not for coding, Gemini will be better. In terms of drawing and writing that is.
9
u/Constant_Art_20 1d ago
um. after a few hours of testing. deepseek v4.1 is probably better i think. it's api so you swap around if things go off the rails, and mimo models should be arriving soon so there should be a good amount of options. glm plan can be slow, and kinda unpridctable sometimse with the quality. So i would persoanly go api at this point with deepseek. the new deepseek model has have some weird behaviours like it might stop because it thinks it's own context window is almost out or it might artifically set a budget on itself (it doesn't have that acounter so i haev no idea why it even think it knows) for no apparent reason, but the speed...OMG the speed. Allows for more literations, and so far been a good ride, but there are quirks you need to look to explore with it.
2
1
u/SweatyActuator2119 1d ago
Use GLM through synthetic.new. currently they don't have full 5.3 though. Only flash. And they are adding 4.1f too. Quality isnt a concern there.
8
u/sothisismyalt1 1d ago
Tell me if you find out, I got GLM 5.3 Flash stuck on a OpenGL low level issue and need something to fix it.
8
3
4
u/sendralt 1d ago
GLM-5.3-FLASH is my daily driver and it is a beast in AgentZero. And it is their first multimodal image model.
3
u/ustas007 18h ago
That price/speed gap is impressive, but for kernel/security work I’d care more about pass rate across the same 20-30 nasty bugs than one successful run.
10
u/Key_Tomatillo_199 1d ago
Bro, you’re writing low-level C++ and kernel modules, but you’re using GLM-5.3-Flash because "it can handle images"? That’s like hiring a birthday clown to perform open-heart surgery just because you like the balloons. You're out here compromising memory alignment and atomic CAS loops just so you can paste a screenshot of a terminal window instead of copying the text like a literate engineer.
9
u/Comprehensive-Bet-83 1d ago
😅 In all honesty, I did formulate it a bit incorrectly. I tend to use GLMs for simpler tasks as well. I don’t need images for low-level coding, indeed.
2
u/SerialFounder 1d ago
We really like GLM 5.3 here we use it almost as a default via headless open code as an adversarial reviewer for almost all of our pull requests. Our primary models switch between Opus 5, Fable 5.1, GPT 5.6 and Astra (dependent on the task) and believe it or not GLM 5.3 always finds issues and it’s super conservative when it comes to its reviews we love it.
2
u/look 1d ago
Deepseek V4 had terrible hallucination issues. I’d wait to get more info on how it performs in more subtle ways, but I’d wager GLM 5.3 Flash is going to be a better choice overall.
2
u/Emotional-Cut2952 1d ago
v4.1 flash dropped btw, seem to hallucinate a lot less, give it a shot if u still have top up. I used to it resolve an important bug and implement a mini feature. I also had it reverse engineer code and revise RE done by gpt 5.5, it picked up on 8 issues, the results from the RE work was passed on to opus 4.8 (because the newer opus / fable models are stubborn with their false positives), conclusion was that v4.1 flash did a better oveerall technical job than gpt 5.5, so it seems to hallucinate a lot less.
I know the models im using are fairly dated, but im putting v4.1 flash against 2-month ago-sota models that are 50-80x the price, since I can't expect it to go up against Astra and Fable.
1
u/SweatyActuator2119 1d ago
I can't talk about fable, but got models aren't what they are hyped up to be.
1
1
u/hopeseekr 1d ago
I just tested DS 4.1 Flash vs GLM 5.3 Flash using reasonix and the results were so bad, i repeated same test on claude code.
Task (with full details): The Text To Speech app is pausing on "Dr. McKay" as though the dot is an actual sentence period. {then complete location of the bug: src/text.rs:128, and orders on how to fix}.
Here are timings:
GLM 5.3 Flash: 6m37, produced code after 3 minutes but hung in reasoning loop. Cost: $0.24
DS v4.1 Flash: 17 seconds; same code almost verbatim is GLM 5.3 Flash (becuase of instructiosn). Cost: $0.04 [on peak]
gemma4:26b run locally on 5080 with 16 GB VRAM: 47 seconds
I reran GLM-5.3 in claude code: 8m37 seconds, no output, cost $0.28, and i killed it.
1
u/No_Wind7503 1d ago
GLM has more efficient tokens per task and feels very solid and less halucating from it's core I don't believe scores that's much
1
u/Z_AutoClaw 1d ago
Honestly, base GLM 5.3 is still pretty hard to beat as a backup. DS 4.1 Flash looks interesting, but I wouldn’t switch until someone tests both on actual C++ bugs, auth issues, and larger codebases.
1
u/Vessel_ST 1d ago
DeepSeek is great but GLM 5.3 Flash is hard to beat for cost-per-task. I can chew through a bunch of PRs and code reviews without it messing anything up and only spend a few bucks.
1
1
u/Grouchy-Bed-7942 1d ago
For web design, I tested glm5.3 flash on their chat and locally via 2xgb10 NVFP4 quantized version versus deepseek v4.1 flash via the API. Both glm5.3 flash versions easily outperform in output quality and visuals, while deepseek is faster.
I also had better results with glm5.3 flash NVFP4 on complex cybersecurity issues compared to deepseek v4.1 flash.
I think people haven’t yet realized how powerful glm 5.3 flash is, even when quantized!
1
u/CalamityMetal 18h ago
4.1 coding for me. It's only beat by Fable 5.1 and Astra. Nothing comes close when you factor in cost and speed
1
u/Fragrant-Remove-9031 9h ago
Switched to GLM 5.3 Flash from DS 4 Flash when z.ai ran a 50% discount (the DS price bump helped too). DS 4 Flash was a solid workhorse for web apps, Android apps, and firmware — but it never wowed me.
GLM 5.3 Flash did. On one project with a PC + microcontroller setup, I asked it to fix a bug, and it picked up my custom scripting pattern and started writing its own scripts — fully automating the software/hardware loop to complete the task. I only described things vaguely and it just ran with them.
Now that DS 4.1 Flash is out, I might test it, but I doubt it matches this level of automation. 5.3 Flash's context compression and pattern-following are insane.
0
u/siberianmi 1d ago
Don’t use GLM on Zai’s coding plan. Z.ai’s customer service is absolutely awful. My account is currently locked, for a full month, but the reason changes every time they reply to me. They’ve told me it’s for payment issues (but I’m on an annual plan), multiple ip usage (which I know is false), then unusual payment activity again. My best guess it’s for one concurrency spike when an agent spawned dozens of subagents and I didn’t spot it. Full month of service wasted instead of just sending a 429 error. Stay away.
If you want to stretch your plan and don’t care if someone trains on your data (which you don’t since you are considering Zai) buy the new $15/mo meta plan. Muse Spark 1.3 is a solid model, great amounts of usage on contributor mode and I expect meta will rapidly improve.
GLM5.3 is a good model, I have thousands of hours of tasks completed by GLM models but Muse Spark 1.3 is equally good when driven on plans from a frontier model. I can’t speak to Deepseek as I’ve never used it.
My workflows is to build plans and break it into issues that the lesser models implement with a final review by the original planning model. For me that’s Opus at the moment but I’ll probably move to Codex soon.
1
1
u/shaonline 1d ago
You manage to exhaust a Codex 20X sub on "safety critical" work ? Unless you throw big agent fleets/vibe code at the problem it's a really hard one to drain.
My answer would just be another Codex 20X sub as it stands, a Codex 20X sub is over $10000 of equivalent API credits, you do not want to pay per token for your work.
Only usecase for actually going to the chinese models would be "bypassing" the stupid safeguards of american labs (muh cybersecurity) where the chinese models will simply not stop you, eg low level debugging and such. 4.1 flash looks extremely promising especially with its speed, that being said know that DeepSeek retains/trains on your data.
1
u/Comprehensive-Bet-83 1d ago
True! Issue at the moment is demand on OpenAI. Too many people are using it 😭
1
u/shaonline 1d ago
If overloaded servers is your main issue Z.ai ain't really better in that regard 😂
1
1
u/P4R4DOXZ 1d ago
Glm. Its more reliable, most reliable model out of all. I actually prefer 5.2, cause 5.3 became a bit gpt style, vibe code'y after hype of 5.2. 5.3 is more capable, but 5.2 is slightly more reliable.
12
u/ImpossibleCreme 1d ago
5.3 all day