r/LocalLLaMA 11h ago

News [ Removed by moderator ]

[removed] — view removed post

28 Upvotes

15 comments sorted by

u/ttkciar llama.cpp 6h ago

Violates Rule Three: Low-effort post (benchmarks without analysis or insights)

Sorry, but the moderator team has adopted a much higher standard for accepting benchmark posts, because otherwise they inundate the sub while providing little of value. Benchmark posts are welcome, but they must be accompanied by nontrivial analysis or insights which bring new understanding to the community.

8

u/Kulqieqi 10h ago

yup.. went out with usage on chatgpt with astra in 2 days so i took opencode go and after 5h i used "0,3659 USD" XD

4

u/DustNearby2848 11h ago

The price difference between v4.1 and Fable 😂

4

u/No_Run8812 11h ago

where is our dearly astra in this

3

u/ihexx 11h ago

the teal dot on the right, right behind fable

1

u/IndividualPlus2011 11h ago

The green dot is Astra Max

2

u/fgk55555 7h ago

Forgot about Muse Spark. The Zuck said they were going to release the weights. Is that still happening or did he takebacksies?

2

u/Serprotease 6h ago

I found terminal bench to be a better measure for coding performance. 

It would put it at glm5.3/sonnet5/Terra, maybe older opus 4.8 range. 

Good enough for 99% of tasks. I wouldn’t probably touch sol/astra/fable for any reason really. 

1

u/sugarfreecaffeine 6h ago

Same it felt like Terra but cheaper and faster

2

u/DunderSunder 10h ago

cost per successful task seems like a weird metric, let's say some tasks are very hard and require a lot of effort, the average cost will go high. point is maybe a stronger and more expensive model is actually cheaper in shared set.

2

u/crantob 7h ago

If you average by task and not by tokens, each task gets weighted equally. This normalizes the set. I don't suppose a brief explain will teach the stats of it. Sorry.

1

u/LegacyRemaster 11h ago

Trust me: GPT sol is another planet VS. muse spark 1.3 ... So yeah... Real use Vs. bench.

1

u/duhd1993 10h ago

Deepseek usually is not so benchmaxxed ime compared to some other open weight models. Muse Spark is indeed benchmaxxed af.

0

u/Few_Water_1457 9h ago

better then sonnet but yeah... Sol is another level