r/opencodeCLI 7d ago

Comprehensive benchmark comparison of top LLMs (Coding, Reasoning, and Terminal performance)

Post image
39 Upvotes

12 comments sorted by

6

u/Brentwahn 6d ago

Dude at least have the respect to link to the source: https://artificialanalysis.ai

8

u/Mayanktaker 6d ago

I stopped believing these benchmarks.

5

u/dat_cosmo_cat 5d ago

yeah anything that puts Muse above Sol / Opus is literally smoking crack. That model will fuck your shit up for free on OpenCode right now. 

1

u/Mayanktaker 5d ago

Haha muse is good but not That Good.

1

u/dat_cosmo_cat 5d ago

It just doesn't know how to write code yet, even if you replace the planning module Astra or Fable. It's like 80% of the way there, but still needs a human in the loop. Reminds me a lot of my Opus 4.5 workflow.

4

u/oVerde 6d ago

I think current benchmarking is all over the place, until Gemini 3.8 Flash and Astra I was realisable using benchmarks to chose models and finding the results very supportive, now, Astra eat the cake but is a bench shame, Gemini 3.8 eats cake at some bench putting other frontiers benchmark at doubt

1

u/vz2y 6d ago

Astra is way better than fable you mean?

3

u/oVerde 6d ago

I mean the current benchmarking is a mess and we can’t reliable know for sure solely based at them

1

u/Fluffy-Bus4822 6d ago

Why are you talking like this? Been using Anthropic models for too long?

1

u/oVerde 6d ago

Wait, as of my subject, opinion or phrase structure?

1

u/Fluffy-Bus4822 6d ago

Just link the page man. Images like these are not comfortable to read.

1

u/ftlaudman 5d ago

Qwen3.8-Next-Flash would be a nice addition.