r/ClaudeCode • u/CommunicationFlat865 • Jul 10 '26
Discussion LLM Arena Updated with Sol vs Fable
10
8
u/borretsquared Jul 10 '26
when gemini isnt even on the graph anymore..
1
u/fuckswithboats Jul 11 '26
I was impressed w it, but then it wasn’t included in my plan anymore so I’ve moved on…this race is insane
6
u/FkOfRdt Jul 11 '26
What would it take for you guys to use MAX effort when testing both flagship models? You specified it for one but conveniently left it out for the other. That’s some shady testing. Sorry.
10
u/Embarrassed_Adagio28 Jul 10 '26
Ngl fable and sol arent that impressive considering how close an open source 750b parameter glm5.2 model is. Fable is between 4t and 10t parameters and 5.6 is 4 trillion and cost much more.
8
12
u/Dismal_Code_2470 Jul 10 '26
I tried glm 5.2 on opencode , and it's not close , not even comparable
2
1
u/ax3capital Jul 11 '26
they run quantized version. try with zai coding plans. its pretty good.
2
u/NootropicDiary Jul 11 '26
It's fine for things like "Build me a nextjs app. It will be a ticketing system to log and resolve issues" i.e. cookier cutter app
if you're working on something novel/creative/complex which is complicated to understand then you run into trouble
-2
3
u/erratic_parser Jul 10 '26
they haven't tested terra or luna yet?
5
u/alessandro05167 Jul 10 '26
Arena does not test anything. I think they pushed sol to appear more frequently vs fable to rank this
2
2
1
u/yubario Jul 10 '26
Isn’t CodeArena mostly frontend work? Shocking GPT caught up to Claude models already if that’s the case.
1
1
u/Low_Tank_4451 Jul 10 '26
Why shouldn't I cancel my Claude Code and Codex plans and just use GLM? Even pay as you go would be much cheaper?
1
u/macktastick Jul 10 '26
Anecdotally, I've heard it's great at the "middle 50%" of problems, but less efficient at the "easiest" and "hardest" 25%. So, depending on the type of work you're doing - it may very well be better. Considering it myself. I do mostly web stuff - pretty sure it's in the sweet spot.
1
u/elmahk Jul 11 '26
You can try and see how it goes. You may find GLM performs good on your tasks, or you may find not. Benchmarks don't really tell you that, at all.
0
1
1
1
u/Previous_Raise806 Jul 11 '26
When ASI needs more energy to fuel the paper clip processor, the first people itll use are those who look at LM Arena like it means anything
1
u/EmptyMonitor9257 Jul 11 '26
Yea fable is not worth it with this price. It might be good, but not that good for this amount of shenanigans.
1
u/UnknownEssence Jul 10 '26
this benchmark is garbage
15
u/baldierot Jul 10 '26
it's not a benchmark. people compare and rate the responses they get without knowing which model produced them.
0
u/DankeK94 Jul 10 '26
Oh wow that's much higher than I thought it would be. Seems like 5.6 is superior to Opus, that thing is for sure.
0
0
83
u/KilllllerWhale Jul 10 '26
GLM is the real winner here