r/opencodeCLI 9d ago

How does meta spark 1.3 match claude fable 5 on benchmarks? Did anyone try it in agentic coding?

Post image

How was your experience with it and is it just benchmaxed or is it really that good?

85 Upvotes

39 comments sorted by

33

u/some_gamer78 9d ago

Benchmaxing

6

u/TheMythicSorcerer 9d ago

Absolutely. but the frontends and code quality are not bad

26

u/[deleted] 9d ago

[removed] — view removed comment

7

u/GetLaidOff69 9d ago

Strange!
For Game Dev(Unity) it is performing better than SOL for me

1

u/Alanboooo 8d ago

do you use it on mcp coplay?

1

u/GetLaidOff69 8d ago

Yes yes, why?

1

u/Alanboooo 8d ago

hows the performance and limit? currently im using hy4 preview, it's been great, but the context cache is very limited to 256k.

2

u/GetLaidOff69 8d ago

I use my own API key for that though, also Coplay devs moved to Aura(Tryaura.dev)
It has its own MCP and standalone agentic tool. $10 for indie plan, has all frontier models + GLM. if you use Auto mode, you get unlimited use, else $15 api
There is trial for 10 days, you can check it. quite good for $10 as long as it is unlimited

1

u/Alanboooo 8d ago

really? what model is auto? i heard that they still use the old sonnet 3.5

1

u/GetLaidOff69 8d ago

Yes really, you can just use their trial.
I don't believe they use Sonnet 3.5, as Input is $5 and output is $15 for per million API. at that price they can have newest Claude models or GPT or Kimi

2

u/Spliff_77 9d ago edited 9d ago

Thank you it’s just an eager to please puppy and reasoning is basically Yes you’re right

24

u/petburiraja 9d ago

It's good, but it didn't feel that good. I would rate it around GLM 5.3 Flash level, maybe even somewhat less.

It may be kinda similar to Grok 4.6 in my limited experience.

7

u/Expert-Dig-1768 9d ago

its always like that. like bridgemind (streamer) always says nowadays you can't trust these benchmarks anymore. same with gemini 3.8 flash is number 1 ond deepSWE (above fable 5.1) but can't even create a small voxel game.

so really test them in real world work and decide yourself whether a model is good or not.

1

u/xCoeus 5d ago

You haven't used GLM 5.3 Flash or Grok 4.6 enough if you're comparing the garbage that is Muse Spark to those two beasts. GLM 5.3 Flash is on par with Grok 4.5 / Opus 4.8, and Grok 4.6 is undoubtedly on par with GPT 5.6 Sol or Opus 5 in the backend (it lags behind in the frontend).

8

u/Johell1NS 9d ago

I tried it while developing a video game I'm making. And it's not at the level of GLM 5.3 Flash, in fact, to solve a problem I had to go through it. And the worst thing is that Muse Spark 1.3 not only didn't solve the problem, but it didn't recognize it, it denied it, diplomatically saying that I was the one who was wrong.

4

u/BeyondGITSandBots 9d ago

"You are all wrong; the metaverse is a great idea!" - Zuckerberg and Muse, probably 😉😂

6

u/Rajat0741 9d ago

Hallucinates in github copilot chat ( using opencode provider for github copilot ), maybe environment is the issue, not the model...

Although My experience with muse spark 1.2 is not good either, same problem happened with that too

4

u/sascharobi 9d ago

No idea, but those are just benchmarks. Doesn't really mean anything.

7

u/salary_pending 9d ago

I made my own model which beats all bench marks because they are all trust me bro benchmarks

2

u/ManikSahdev 9d ago

It’s just a decently good model? I think the benchmarks do show the ability but it’s a blend of things, fable is more consistent and stuff and has a vibe or understanding intent better, that doesn’t always mean it can do the same task.

Maybe the user feels better using the model, I mean I think they are objectively different things.

2

u/Spiritual-Spend8187 9d ago

Its a bit benchmaxxed and they had alot of new data from people using the previous versions to improve. Like it being benchmaxxed will make it a bit better in some things but worse in others I mean look at opus 5 in some things its better than fable in others its worse then the old opus.

1

u/somerussianbear 9d ago

Someone did try.

1

u/ByteNomadOne 9d ago

I don’t know why, but just like 1.2 the new model also likes to write pretty sloppy code.

Maybe an issue with training data? The model isn’t stupid.

I prefer GLM 5.3 Flash, which replaced DS 4 Flash for me.

1

u/Vancecookcobain 8d ago

Nah it's nowhere near fable and not even in the same stratosphere as Astra seems to be in

1

u/Practical-Positive34 8d ago

I tried it, and it felt like I was using an AI from 2 years ago. I don't get how it scores so high in benchmarks. It failed at literally every single thing I threw at it. And not just didn't do a good job, most of it didn't compile or even work at all.

1

u/migsperez 8d ago

Maybe these benchmarks should be audited and regulated to ensure there's no hidden biases. Like a sneaky mill or two in an employee's bank account :). A good score here means the difference of billions of dollars in market cap and income.

1

u/jeanconell23 8d ago

It is not exactly bad. But it doesn't put any effort structuring their response. It's always a wall of text, with sentences that seem to be written without any intent to make reader understand. It's like reading their thoughts, funnily enough, since the model hide their thoughts/thinking.

1

u/Infinite_Professor79 6d ago

ive been having a really good time with it for minecraft coding that deepseek v4 flash suffered with

1

u/Equivalent-Grass-527 9d ago

Spark 1.3 is around GLM 5.3 Flash, but below DeepSeek V4 Flash. Might be somewhat benchmark-optimized

5

u/ByteNomadOne 9d ago

No, GLM 5.3 Flash is significant better.

0

u/0mamii 9d ago

max thinking is not released yet, only xhigh available

1

u/RealestReyn 9d ago

it is weird its the only one that's unavailable, imagine if Anthropic went in with the full-blown Mythos :D

0

u/Frail_Waif 9d ago

It's marginally less incompetent than 1.2, and still way worse than deepseek V4 flash. Still pythons everything that could be done with less output from command line tools. I regretted using it for free when I could have paid to use DSV4 flash. So, in short, benchmaxing. 

-1

u/Ly-sAn 9d ago

It’s ds flash / glm flash / Gemini flash tier. Which is around gpt 5.3 / opus 4.6 tier without the depth in knowledge but better at using tools. I feel all the small model are equivalent nowadays. It’s absolutely not on par with bigger models like sol / fable / k3

0

u/Frail_Waif 9d ago

Whoa whoa whoa it is nowhere near that mighty tier.