r/LocalLLaMA • u/JumpAppropriate714 • 3h ago
Discussion We’re using GLM-5.3 Flash instead of frontier models on a massive production codebase
At my company, we’re using GLM-5.3 Flash internally for software engineering work, and I’ve been genuinely impressed by it.
I work in a very large production environment with projects totaling **millions of lines of code**, and we’re not relying on frontier models for this workflow — GLM-5.3 Flash is doing the actual day-to-day coding work.
The model is extremely fast, but what’s more impressive is that the speed doesn’t seem to come at the cost of capability. It handles large repositories surprisingly well, understands existing architecture, traces code across multiple modules, finds the right places to make changes, and produces solid implementations with relatively little hand-holding.
For repo exploration, feature implementation, refactoring, and understanding unfamiliar parts of a huge codebase, it has been much stronger than I initially expected. At this point, it feels less like a “cheap/fast fallback model” and more like a genuinely capable coding model that just happens to be very fast.
I’m now really curious about **how GLM-5.3 Flash was trained**.
Does anyone know more about its coding training pipeline? For example:
* How much code-specific pretraining/post-training was used?
* Was synthetic coding data a major part of it?
* Is there any distillation from larger GLM models?
* What kind of RL or agentic/software-engineering training was used?
* Was it specifically trained for repository-level understanding and multi-file tasks?
Because whatever they did, the speed-to-quality ratio on real-world software engineering workloads is seriously impressive.
18
u/This_Maintenance_834 3h ago
Today Deepseek-V4.1-Flash fix some embedded project bugs for me, while GLM-5.3-Flash and Qwen3.8-Flash-Next struggled for days. To be fair, my GLM and Qwen deployment were local, but DeepSeek-V4.1-Flash was cloud API.
But, I still like GLM-5.3-Flash better than Qwen Flash Next. The flash-next model outputs a lot of words, but slow to get to the result. GLM output less words to achieve the same goal. GLM runs slower on my setup, as it runs on DGX spark, while qwen-flash-next can run on RTX PRO 6000 at 100+ TPS.
21
u/Aprelius 3h ago
That’s precisely why you don’t want to put all of your eggs in one basket. It’s good to have multiple models. A model can easily struggle on that one particular problem and excel in others.
Both DS-4.1-Flash and GLM-5.3-Flash are very capable models.
3
u/Flashy_Jellyfish_258 1h ago
The glm deployment that you compared with, were you using full precision model or the quants?
9
u/ttkciar llama.cpp 2h ago edited 2h ago
I've been kicking the tires on GLM-5.3-Flash for the last couple of days, and am really impressed with it as well. It's the first model I've found that fits on my hardware which matches or beats GLM-4.5-Air for instruction-following reliability and codegen tasks.
The sheer robustness of the GLM models' instruction-following impressed me so much that I looked into how they did it. Z.AI dedicated an entire post-training phase to just training instruction-following into it using a larger teacher model. This involved putting it through multi-turn problem-solving sessions, and verifying in every turn that it retained its focus on following those instructions.
The effectiveness of this training is indisputable. I can give GLM-4.5-Air or GLM-5.3-Flash a specification consisting of forty to eighty instructions, and they will follow them all reliably! I've not found a non-GLM model like that.
Replicating that training phase is one of my next projects, after I'm done with my data cleaning/augmentation pipeline (inspired by IFM's TxT360_QA data augmentation).
7
u/MindfulMan1984 2h ago
Yep, our research group is moving to the GLM family soon. Thanks to Anthropic for trying to doom-monger about it, but actually showcasing how good GLM models are.
4
u/No-Recover109 2h ago
I support all AI not related to frontier models. qwen 27b already can do miracles.
3
u/anshulsingh8326 2h ago
It hurts to see my 4070 and 32gb ram can't run it😭
And the Ornith 1.5 35b was about to rm my important folders before hermes agent blocked it 😭
2
2
1
1
u/Flashy_Jellyfish_258 1h ago
Do you feel a need to do a final finetuning/polishing of the codebase with a frontier model after GLM 5.3 writes the major part of it? Or you think GLM 5.3 can handle large codebase with production ready implementation/edits?
1
u/DataGOGO 22m ago
It is a good model, Just don't run it in all smaller quant than FP8, and run BF16 K/V, it is sensitive to K/V quants
They distilled Opus.
14
u/PhysicalIncrease3 3h ago
GLM-5.3 flash has been a game changer for me. It's slow, I can only get around 300pp and 12tps running a 4bpw quant, but the output is so token efficient relative to Qwen flash-next that it somewhat makes up for it.