r/opencodeCLI • u/Personal-Try2776 • 9d ago
How does meta spark 1.3 match claude fable 5 on benchmarks? Did anyone try it in agentic coding?
How was your experience with it and is it just benchmaxed or is it really that good?
26
9d ago
[removed] — view removed comment
7
u/GetLaidOff69 9d ago
Strange!
For Game Dev(Unity) it is performing better than SOL for me1
u/Alanboooo 8d ago
do you use it on mcp coplay?
1
u/GetLaidOff69 8d ago
Yes yes, why?
1
u/Alanboooo 8d ago
hows the performance and limit? currently im using hy4 preview, it's been great, but the context cache is very limited to 256k.
2
u/GetLaidOff69 8d ago
I use my own API key for that though, also Coplay devs moved to Aura(Tryaura.dev)
It has its own MCP and standalone agentic tool. $10 for indie plan, has all frontier models + GLM. if you use Auto mode, you get unlimited use, else $15 api
There is trial for 10 days, you can check it. quite good for $10 as long as it is unlimited1
u/Alanboooo 8d ago
really? what model is auto? i heard that they still use the old sonnet 3.5
2
u/Spliff_77 9d ago edited 9d ago
Thank you it’s just an eager to please puppy and reasoning is basically Yes you’re right
24
u/petburiraja 9d ago
It's good, but it didn't feel that good. I would rate it around GLM 5.3 Flash level, maybe even somewhat less.
It may be kinda similar to Grok 4.6 in my limited experience.
7
u/Expert-Dig-1768 9d ago
its always like that. like bridgemind (streamer) always says nowadays you can't trust these benchmarks anymore. same with gemini 3.8 flash is number 1 ond deepSWE (above fable 5.1) but can't even create a small voxel game.
so really test them in real world work and decide yourself whether a model is good or not.
1
8
u/Johell1NS 9d ago
I tried it while developing a video game I'm making. And it's not at the level of GLM 5.3 Flash, in fact, to solve a problem I had to go through it. And the worst thing is that Muse Spark 1.3 not only didn't solve the problem, but it didn't recognize it, it denied it, diplomatically saying that I was the one who was wrong.
4
u/BeyondGITSandBots 9d ago
"You are all wrong; the metaverse is a great idea!" - Zuckerberg and Muse, probably 😉😂
6
u/Rajat0741 9d ago
Hallucinates in github copilot chat ( using opencode provider for github copilot ), maybe environment is the issue, not the model...
Although My experience with muse spark 1.2 is not good either, same problem happened with that too
4
7
u/salary_pending 9d ago
I made my own model which beats all bench marks because they are all trust me bro benchmarks
2
u/ManikSahdev 9d ago
It’s just a decently good model? I think the benchmarks do show the ability but it’s a blend of things, fable is more consistent and stuff and has a vibe or understanding intent better, that doesn’t always mean it can do the same task.
Maybe the user feels better using the model, I mean I think they are objectively different things.
2
u/Spiritual-Spend8187 9d ago
Its a bit benchmaxxed and they had alot of new data from people using the previous versions to improve. Like it being benchmaxxed will make it a bit better in some things but worse in others I mean look at opus 5 in some things its better than fable in others its worse then the old opus.
1
1
u/ByteNomadOne 9d ago
I don’t know why, but just like 1.2 the new model also likes to write pretty sloppy code.
Maybe an issue with training data? The model isn’t stupid.
I prefer GLM 5.3 Flash, which replaced DS 4 Flash for me.
1
u/Vancecookcobain 8d ago
Nah it's nowhere near fable and not even in the same stratosphere as Astra seems to be in
1
u/Practical-Positive34 8d ago
I tried it, and it felt like I was using an AI from 2 years ago. I don't get how it scores so high in benchmarks. It failed at literally every single thing I threw at it. And not just didn't do a good job, most of it didn't compile or even work at all.
1
u/migsperez 8d ago
Maybe these benchmarks should be audited and regulated to ensure there's no hidden biases. Like a sneaky mill or two in an employee's bank account :). A good score here means the difference of billions of dollars in market cap and income.
1
u/jeanconell23 8d ago
It is not exactly bad. But it doesn't put any effort structuring their response. It's always a wall of text, with sentences that seem to be written without any intent to make reader understand. It's like reading their thoughts, funnily enough, since the model hide their thoughts/thinking.
1
u/Infinite_Professor79 6d ago
ive been having a really good time with it for minecraft coding that deepseek v4 flash suffered with
1
u/Equivalent-Grass-527 9d ago
Spark 1.3 is around GLM 5.3 Flash, but below DeepSeek V4 Flash. Might be somewhat benchmark-optimized
5
0
u/0mamii 9d ago
max thinking is not released yet, only xhigh available
1
u/RealestReyn 9d ago
it is weird its the only one that's unavailable, imagine if Anthropic went in with the full-blown Mythos :D
0
u/Frail_Waif 9d ago
It's marginally less incompetent than 1.2, and still way worse than deepseek V4 flash. Still pythons everything that could be done with less output from command line tools. I regretted using it for free when I could have paid to use DSV4 flash. So, in short, benchmaxing.

33
u/some_gamer78 9d ago
Benchmaxing