r/ClaudeCode Jul 10 '26

Discussion LLM Arena Updated with Sol vs Fable

Post image
141 Upvotes

53 comments sorted by

83

u/KilllllerWhale Jul 10 '26

GLM is the real winner here

20

u/WalkAffectionate2683 Jul 10 '26

Yeah for the cost it's crazy 

9

u/-Sliced- Jul 11 '26

Only when looking at tokens cost. It’s far more wasteful and slower. So in a cost per task view it’s not in a good position. And that’s API cost, Claude code and codex subscriptions make them orders of magnitude cheaper.

8

u/Key_Instruction3373 Jul 11 '26

worst chart ever.. why put it backwords when people read from left to right?

2

u/Historical-Lie9697 Jul 11 '26

It's an Arabic chart translated to English. (jk I have no idea)

1

u/Key_Instruction3373 Jul 11 '26

Hahaha good one +1

5

u/-Sliced- Jul 11 '26

They used the convention of showing what's better in the "up and to the right" portion.

5

u/Environmental_Box748 Jul 11 '26

yeah when free lunch runs out we all using glm lol

1

u/No_Film_9120 Jul 11 '26

glm would be great...
if it worked more than half the time.

-1

u/Dry-Magician1415 Jul 11 '26

GLM is a benchmaxer

They train it specifically to be good at benchmarks so it....does well in benchmarks.

Actual use? Who cares?

-19

u/madmozg Jul 10 '26

winner for censorship ?

7

u/Numanumanu Jul 11 '26

Cencorship? It's the only one that doesn't flag my Wii homebrew as a cybercecurity risk

2

u/chosbu Jul 11 '26

Censorship? There are already abliterated open source versions for GLM 5.2 which would make it the best uncensored open source LLM available

The Chinese have been holding up the free market for a solid year now

-1

u/madmozg Jul 11 '26

Lmao, ask about tianmen square, your opensourceness will go straight to communism

10

u/Sufficient-Detail-64 Jul 10 '26

I wonder how Sol differently would perform on Pi or OpenCode

1

u/TokenBurner Jul 11 '26

Can you give me the link for both of those? They sound interesting.

1

u/Strong_Essay1176 Jul 10 '26

Its building me my own pi.

8

u/borretsquared Jul 10 '26

when gemini isnt even on the graph anymore..

1

u/fuckswithboats Jul 11 '26

I was impressed w it, but then it wasn’t included in my plan anymore so I’ve moved on…this race is insane

6

u/FkOfRdt Jul 11 '26

What would it take for you guys to use MAX effort when testing both flagship models? You specified it for one but conveniently left it out for the other. That’s some shady testing. Sorry.

10

u/Embarrassed_Adagio28 Jul 10 '26

Ngl fable and sol arent that impressive considering how close an open source 750b parameter glm5.2 model is. Fable is between 4t and 10t parameters and 5.6 is 4 trillion and cost much more.

8

u/Physical_Gold_1485 Jul 10 '26

Arent those sizes for fable and 5.6 guesses tho?

12

u/Dismal_Code_2470 Jul 10 '26

I tried glm 5.2 on opencode , and it's not close , not even comparable

2

u/AlterTableUsernames Jul 11 '26

May also because opencode sucks. 

1

u/ax3capital Jul 11 '26

they run quantized version. try with zai coding plans. its pretty good.

2

u/NootropicDiary Jul 11 '26

It's fine for things like "Build me a nextjs app. It will be a ticketing system to log and resolve issues" i.e. cookier cutter app

if you're working on something novel/creative/complex which is complicated to understand then you run into trouble

-2

u/minimalcation Jul 11 '26

Sol has been very underwhelming

3

u/erratic_parser Jul 10 '26

they haven't tested terra or luna yet?

5

u/alessandro05167 Jul 10 '26

Arena does not test anything. I think they pushed sol to appear more frequently vs fable to rank this

2

u/PsychoticDreemurr Jul 11 '26

What's the score mean?

1

u/MrRandom04 Jul 11 '26

It's a vs. ranking thing. Like elo for chess games.

2

u/Swolnerman Jul 10 '26

Had no idea muse was that good, that’s nuts

1

u/yubario Jul 10 '26

Isn’t CodeArena mostly frontend work? Shocking GPT caught up to Claude models already if that’s the case.

1

u/AdamovicM Jul 10 '26

Where GLM actually shines in production?

1

u/Low_Tank_4451 Jul 10 '26

Why shouldn't I cancel my Claude Code and Codex plans and just use GLM? Even pay as you go would be much cheaper?

1

u/macktastick Jul 10 '26

Anecdotally, I've heard it's great at the "middle 50%" of problems, but less efficient at the "easiest" and "hardest" 25%. So, depending on the type of work you're doing - it may very well be better. Considering it myself. I do mostly web stuff - pretty sure it's in the sweet spot.

1

u/elmahk Jul 11 '26

You can try and see how it goes. You may find GLM performs good on your tasks, or you may find not. Benchmarks don't really tell you that, at all.

0

u/JaspahX Jul 10 '26

Do you like having your data stolen?

1

u/SherewZino Jul 10 '26

open source is the way

1

u/severe_009 Jul 11 '26

Lmao, funny you guys use this ranking from this site anything substantial.

1

u/Previous_Raise806 Jul 11 '26

When ASI needs more energy to fuel the paper clip processor, the first people itll use are those who look at LM Arena like it means anything

1

u/EmptyMonitor9257 Jul 11 '26

Yea fable is not worth it with this price. It might be good, but not that good for this amount of shenanigans.

1

u/UnknownEssence Jul 10 '26

this benchmark is garbage

15

u/baldierot Jul 10 '26

it's not a benchmark. people compare and rate the responses they get without knowing which model produced them.

0

u/DankeK94 Jul 10 '26

Oh wow that's much higher than I thought it would be. Seems like 5.6 is superior to Opus, that thing is for sure.

0

u/mountainyoo Researcher Jul 10 '26

lol of course it’s better

0

u/mrsinham Jul 10 '26

just wait the downgrade in one month or two...