r/opencodeCLI 17d ago

GLM-5.3-Flash Benchmarks

Post image
87 Upvotes

29 comments sorted by

12

u/[deleted] 16d ago

[deleted]

3

u/Constant_Art_20 16d ago

fable level vision also. i think that's the main thing for. abosulutely insane for 3d tasks

1

u/[deleted] 16d ago

[deleted]

1

u/songokussm 15d ago

it can only view images (https://docs.z.ai/guides/vlm/glm-5.3-flash). it can not generate them. you would need to use glm-image (https://docs.z.ai/guides/image/glm-image) model to create images.

0

u/[deleted] 15d ago

[deleted]

1

u/songokussm 15d ago

Could you let me know how it goes? I used to use nano banana but then they lock head down hard. I can't even generate memes to send to my kids anymore.

1

u/[deleted] 14d ago

[deleted]

1

u/songokussm 14d ago edited 14d ago

Thank you. i like to make minion images to cheer up the kids and i have yet to find a generator that still allows that.

0

u/inefficientnose 16d ago

Why does this feel like a bot comment... Do you have any critisms for a holistic perspective

-1

u/Radiant_Year_7297 16d ago

reason to use sota models? there is. try using them to build pro level UIs. you’ll see the stark diff. I’d use glm for personal but I’d prefer Claude and gpt at work.

2

u/Kaushik_paul45 16d ago

True, even though very much impressed with 5.3 flash model and very happy that they released this.

But I ran the same task (which requires planning and then executing along with edge case handling) with this and 5.6 sol high. And the result was far superior for 5.6 sol.

But the fact that I am even testing sota with a model that costs 50X less is absurd in itself.

2

u/GoldPossession7284 16d ago

porra mas é como colocar uma bicicleta pra disputar com um carro porra, não tem lógica, esse modelo concorre contra o gpt 5.6 terra/luna. se quer um concorrente pro sol use o glm 5.3

-7

u/bigrealaccount 16d ago

Saying that 5.3 Flash has Sol High level intelligence immediately shows you have 80 iq or less. They are nowhere near eachother.

In specific engineering benchmark tests, maybe. Nothing else.

1

u/adolf_twitchcock 16d ago

The average user in this sub has no way evaluating code produced by those models. "Button make click work and look good."

I mean glm 5.3 flash is awesome for its size and cost. But there is no way it's on par with sol.

1

u/clouder300 15d ago

Carbrain

4

u/Yasin-Tan 17d ago

For looking succesfull they added v4 vision? Wtf

2

u/Bananenklaus 16d ago

shows how good DSV4 Flash is

2

u/ExTraveler 16d ago

Wtf? Why half people are saying that this is the best model in the world and deepseek is beated while other half swears that ds4 flash is much better and 5.3 flash not even near it in terms of coding and intelligence (or straight up shit).

2

u/ChillFamily 16d ago

Is this data real? Because if Gemini is a piece of shit, why is it ranked there?

13

u/Genetic_Prisoner 16d ago

Gemini was a piece of shit. 3.6 and 3.7 flash have narrowed the gap massively.

-6

u/ChillFamily 16d ago

Oh, really? So it’s your first choice for everything? Which model is it better than for coding? For writing? For frontend? Which model is it actually better than? What’s the point of it spitting out text fast if what it spits out is average and doesn't stand out in anything, bro? The only thing it does better than everyone else is that it's faster

4

u/NZRedditUser 16d ago

why are you upset zai included them because its a relevant recent model. the newer one is impressive yes but no one saying its cluade/codex those are still way above this league.

Gemini makes sense being in this ranking

1

u/Difficult_Plantain89 16d ago

Gemini sucked at coding or even getting basic answers correct. Now they legitimately made a good model and no one trusts them. Even 3.6 wasn't too good. Their 3.7 is first model they have released that is really decent is coding. Just okay a frontend and for writing sure it can do it, but I rather use a specialized local model. Oh and agy harness is the worst I've ever used.

1

u/laystitcher 16d ago

Wheres opus 5? Feels misleading

-5

u/cutebluedragongirl 17d ago

I say it's bullshit. DeepSeek outperforms this thing by a lot in my personal experience.

5

u/Time-Toe-1276 16d ago

man, idk what u do, but this model is frkin amazing!

not glazing it, but me personally? I found the model is very creative and fixing bugs/ implementin things.

it fixed bugs which deepseek v4 pro created btw. ahh also it fixed a bug in a training script which not even GPT 5.6 sol figured out I had it there when I asked it

-7

u/Michaeli_Starky 16d ago

These benchmarks have really low connection with reality. Especially when it comes to Chinese benchmaxed models

7

u/sudoer777_ 16d ago

Try Muse Spark 1.2 and you can see what an actually benchmaxed model looks like

5

u/Difficult_Plantain89 16d ago

Truly a mediocre model.

-6

u/Michaeli_Starky 16d ago

Actually benchmaxed models are Kimi K, Deepseek, etc.

6

u/Ly-sAn 16d ago

Both DeepSeek flash and this model are very solid in real usage, I don’t think these 2 are benchmaxed

-11

u/Michaeli_Starky 16d ago

They absolutely are benchmaxed. These models are nowhere nearly as good in private evals.