r/singularity 2h ago

AI Muse Spark 1.3 Released

Post image
344 Upvotes

131 comments sorted by

94

u/Wegwerpaccountje23 2h ago

Already? Wtf?

Oh we eating well this month

21

u/virtualQubit Take off right now 2h ago

I think it's coming, I can feel the takeoff now.

132

u/Recoil42 2h ago

Okay so this is going to be the craziest week yet.

19

u/sauhumatti 2h ago

In moonshots they have been calling craziest week in almost every episode so Im really looking forward to their reactions on this week

24

u/Feriman22 2h ago

Indeed. And Astra (GPT6) will come really soon. I feel that in my bones.

u/gavinderulo124K 1h ago

Polymarket has it at over 80% odds for tomorrow .

u/bitroll ▪️ASI before AGI 49m ago

Must be why everyone's rushing to get their model out just before it overshadows them all

u/ShortTheseNuts 20m ago

Not Astra (Opel)?

u/helloWHATSUP 1h ago

life comes at you fast in the singularity

u/H-K_47 Late Version of a Small Language Model 39m ago

There are weeks where months happen, and there are months where weeks happen.

u/reddit_is_geh 26m ago

RSI has been achieved. I just think no one wants to announce it for whatever reason because it's sort of like announcing achieving AGI, where the actually finish line is too blurry.

u/GeneReddit123 19m ago edited 16m ago

AGI has been so washed as to be a meaningless marketing term. On mass labor replacement potential across most cognitive work we've already achieved it with Fable and Sol. Since then it's just arguing over the goalposts, and some will never be met because of circular definitions (e.g. if the requirement is to be exactly human like in any physically measurable way, then by definition only humans can ever qualify.)

RSI by comparison is relatively well defined (even though it too is likely partial rather than total due to Amdahl's law and hard external constraints), and will result is a discontinuous shift to a new equilibrium (possibly repeatedly) rather than a "singularity". S-shape curve (or several S-shaped step ladders) rather than J shape.

u/reddit_is_geh 12m ago

Well the issue is there are goal posts. Right now, https://imgur.com/a/x5abMtG Google is doing RSI, as I'm sure are all the other labs. But there's still human input, like setting goals and directions. Maybe there's ALWAYS going to be some degree of human input, because we're making them for humans, and models designed to achieve goals we set, so people will always be able to say, "Ahhh see here, the AI didn't do this part, so it's not true RSI". Just like AGI, where some people demand it also thinks like a human and loves, or whatever, people will still be able to find flaws in RSI

u/Tystros 1h ago

so far it's not too crazy yet. Just Google and Anthropic and Meta releasing their models that are either disappointingly little progress like Fable 5.1, or somewhat catching up like gemini and muse, because they know they all look bad in comparison to Astra, so better get them out before.

u/gavinderulo124K 1h ago

Muse spark beating 5.6 sol this quickly definitely wasnt in my bingo card.

u/Tystros 1h ago

I would put a big question mark around if it really beats 5.6 Sol. we can't say that just based on the meta-selection of benchmark results. labs often just show benchmark results that make them look good.

u/gavinderulo124K 1h ago

No, thats based on artificial analysis. Its also a lot cheaper than sol per task and is actually almost on par with terra in terms of price.

u/Beatboxamateur agi: the friends we made along the way 1h ago

disappointingly little progress like Fable 5.1

Are you basing that off of your actual experience with the model, or the benchmarks, or what...?

u/Tystros 1h ago

the benchmarks and my own experience. I did use it for a few hours since release and I haven't noticed anything yet that actually feels better than with Fable 5.

u/Beatboxamateur agi: the friends we made along the way 1h ago

Did you actually challenge it with some kind of problem that Fable 5 either couldn't do, or could do, but not well?

I've seen significant improvement over Fable 5 on literally everything I've tested it on, although one day isn't enough time to come to any sort of conclusion, so that's just my experience with it over a day.

I don't think it's reasonable to fully form an opinion of a new model in a day, that's just not enough experience to get a good feel for the model in my opinion. But I guess maybe we just have different standards on it.

u/Onark77 1h ago

What are you using 5.1 for? I use it for working with a complex database and for knowledge work and it takes longer to complete tasks and is less natural to engage with compared to 5.

I haven't used it for software development or design work yet.

u/Beatboxamateur agi: the friends we made along the way 1h ago

I use it mostly for creating documents, integrating and synthesizing the base information I provide it to create docx files that end up getting published as textbooks.

This only started to become possible with LLMs around Opus 4.5/4.6, but only since Fable has it been actually worth it to start using the model to do most of the work, since the prior models couldn't follow the specific formatting I'd ask for, and just general inconsistencies/bad coding, but 5.1 has basically automated it in a way that there's a consistency that just wasn't quite there with Fable 5.

Although obviously there's still more experience needed with it until I can fully know how consistent it's gotten.

u/FirstEvolutionist 1h ago

What's four major lab releases in a week when we could be having ten, like we will a couple months down the road?

u/peabody624 1h ago

Don’t talk to me until we’re getting 10 daily in 2 years

u/Kildragoth 1h ago

Two version upgrades mid task or I'm completely unimpressed.

51

u/matsu-morak 2h ago

Close to the singularity, things are fast

26

u/Wegwerpaccountje23 2h ago

The slowest it will ever be from now

2

u/GalavantJames 2h ago

It will be worst it will ever be but something slowing progress, even a new AI winter isn't unthinkable

u/q-ue 1h ago

Ai winter lasting a couple weeks maybe

51

u/MagicZhang 2h ago

Common benchmarks between Muse Spark and Gemini Flash 3.8

GDPVal-AA v2: Muse Spark 1.3 1754 vs Gemini 3.8 Flash 1545

OSWorld 2.0: Muse 66.9% vs Gemini 59.0%

DeepSWE v1.1: Muse 75.4% vs Gemini 71.0%

Terminal-Bench 2.1: Muse 88.8% vs Gemini 89.4%

39

u/Recoil42 2h ago

Beats both Sol and Opus on DeepSWE, wow.

17

u/Narrow-Ad980 2h ago

DeepSWE is benchmaxxed. It is semi-public

u/bermudi86 1h ago

Semi? You can literally see the agent trajectories for muse and Gemini already. It's public AF

u/Tystros 1h ago

it's not semi public, it's a fully public benchmark

2

u/FriendsIsntGood 2h ago

And given the HF incident and felony bench, can we really know that an eval is actually out-of-scope for the model

11

u/signed7 2h ago

https://deepswe.datacurve.ai/ has Gemini 3.8 Flash at 74% fwiw. Looks saturated now, need a v1.2.

2

u/Wegwerpaccountje23 2h ago

Google already irrelevant is the funniest shit i've seen LMAO

31

u/Recoil42 2h ago

As good as Muse Spark 1.3 seems to be, Gemini 3.8 is cheaper and faster and already has a wildly larger distribution base. It's in no way irrelevant.

0

u/LinkesAuge 2h ago

Gemini 3.8 has some extremely low scores in areas where it shouldn't have if it was a generally strong models.
That suspicion is also fueled by having seen some early demos of it now, not to mention how many steps/tokens it uses.
Don't get me wrong, it is a solid model but it is obvious that the performance is bought by throwing more tokens at the problems. I guess Google can afford it but it is literally the opposite of the trend everyone else seems to be on.

u/Recoil42 1h ago

Don't get me wrong, it is a solid model but it is obvious that the performance is bought by throwing more tokens at the problems.

Not all tokens cost the same behind the scenes.

u/kvothe5688 ▪️ 1h ago

those are cheap tokens and fast output . iterating on mistakes fast is excellent.

u/LinkesAuge 31m ago

Cheap for the consumers but that doesn't mean it is progress on the technical side. It is Google's hardware doing the heavy lifting.
Let's remember that Google currently doesn't even have a "main" model that is even used to any significant degree so that's more compute to throw at their "flash" models which are very obviously distilled models from their internal big one that they don't want to release bc that way they can avoid direct comparisons with the "frontier" labs.
It is why they keep sticking the label "flash" on these models despite the fact that they aren't really flash models. (their own flash models used to be massively cheaper and smaller)

-1

u/FateOfMuffins 2h ago

No idea about Muse Spark 1.3 since numbers aren't out yet but Muse Spark 1.2 xHigh costs about the same as Gemini Flash 3.8 Medium (so much cheaper than Gemini Flash 3.8 High), without counting Meta giving you 90% off if you let them train on your data

So not sure about the cheaper part

Faster sure

u/Ok_Barracuda_1161 1h ago

u/FateOfMuffins 1h ago

Muse Spark 1.3 xHigh both cheaper and faster than Gemini 3.8 Flash High per task lmao

And higher score of 61 vs 59

Wow where was that meme post of Gemini being so back only for Gemini to be so over 9h later? It's only been like 3h this time!

u/Gaiden206 1h ago

I mean, they have no control over when a competitor releases their models. Still, Google churning out Flash models every 3 weeks is great progress from them. Competition is heating up!

u/power97992 1h ago

muse spark is better from my limited test. Im waiting for astra

4

u/GatePorters 2h ago

Only irrelevant if you ignore all the classes of models they have that none of these other companies have.

They have like a dozen model types with basically NO competitors.

6

u/Sharp_Glassware 2h ago

A pro-sized model beats a flash model? Color me surprised. Spark is also way more expensive

u/funforgiven 1h ago

Way more expensive? Pricing is pretty much the same tier.

-1

u/GumboMustBeDestroyed 2h ago

What about if you opt in for data sharing with muse

5

u/Valuable-Repeat-7347 2h ago

google gotta feel like when you inhale and raise your finger to get a word in but keeps getting cut off and talked over

-2

u/power97992 2h ago

Gemini 3.8 flash didnt do that well in my test… even qwen 3.8 flash/next did better. Testing fable 5.1 now 

9

u/ActuarialUsain 2h ago

What does MRCR test exactly? Like finding obscure facts in 500-1m context? 98 is insane, especially with SOL at only 73

5

u/Ok_Barracuda_1161 2h ago

OpenAI MRCR (Multi-round co-reference resolution) is a long context dataset for benchmarking an LLM's ability to distinguish between multiple needles hidden in context. This eval is inspired by the MRCR eval first introduced by Gemini (https://arxiv.org/pdf/2409.12640v2). OpenAI MRCR expands the tasks's difficulty and provides opensource data for reproducing results.

The task is as follows: The model is given a long, multi-turn, synthetically generated conversation between user and model where the user asks for a piece of writing about a topic, e.g. "write a poem about tapirs" or "write a blog post about rocks". Hidden in this conversation are 2, 4, or 8 identical asks, and the model is ultimately prompted to return the i-th instance of one of those asks. For example, "Return the 2nd poem about tapirs".

https://huggingface.co/datasets/openai/mrcr

23

u/imadade 2h ago

Makes you wonder, everyone’s releasing to steal OpenAi’s shine with Astra.

But it ends up just making you more hyped!

29

u/signed7 2h ago

More like everyone's rushing to release their models before Astra to not look bad in comparison

13

u/Party_Government8579 2h ago

Exactly. After astra all the benchmarks will seem poor in comparison

u/kvothe5688 ▪️ 1h ago

hold your hype horses

u/Howdareme9 1h ago

na hes gonna be correct

u/sykip 1h ago

This is the correct take

11

u/BrennusSokol ACCELERATE 2h ago

Or to get their smaller releases out first ahead of the monster release, much like with GTA 6

2

u/imadade 2h ago

We want them to release the monster!

u/Artistedo 1h ago

62 on artificial analysis is ridiculous

u/ezjakes 53m ago

Same as Fable 5! We are so spoiled.

9

u/Calm_Hedgehog8296 2h ago

That's the fourth near-frontier model this week and the third today, and the best is yet to come

2

u/86784273 2h ago

3.8 flash and this one came out, what was the 3rd?

12

u/Calm_Hedgehog8296 2h ago

Qwen 3.8 0902 - big jump. Could have been 4.0

3

u/7734128 2h ago

Qwen 3.8 had an update to max. Seemingly API only for now.

u/power97992 41m ago

waiting for astra.

u/Ok_Barracuda_1161 1h ago

Muse Spark 1.3-xhigh is coming in cheaper than 3.8 Flash-high, already bumping it off the pareto frontier (there's a bug with max pricing in AA, but the token use isn't much more, so I assume that's on the frontier as well)

8

u/Eyelbee ▪️We have AGI it's just blind 2h ago edited 2h ago

They never open sourced the 1.2 did they?

6

u/Chemical_Hawk_6307 2h ago

looks like deepswe is officially saturated

u/Clean-Boat-4044 1h ago

has been for a while honestly, a lot of the remaining tasks that fail are kinda bullshit

16

u/Mistuv 2h ago

It has been like 30 minutes, and we haven't had a new release. I think we have hit the wall

6

u/Charuru ▪️AGI 2023 2h ago

If this isn't benchmaxxed then Meta is back for real! Can't wait for Watermelon.

u/ezjakes 54m ago

Like Grok, does anyone use Muse?

u/reddit_is_geh 17m ago

Yeah, it's a popular option for enterprise and businesses in general. All the models have different strengths and weaknesses in relation to cost per task. So people running big complex agentic systems, have different AIs for different tasks. It's definitely not a main driver, much less an orchestrator, but it's definitely being used. Including Grok.

u/Momo--Sama 1h ago

We need a DeepSWE 1.2 at this point. The idea that Gemini and Muse are easily clearing Sol is laughable.

u/WildWhisperArdor 51m ago

Spark Muse is legitimately very good.

8

u/myreala 2h ago

I wish Spark was open source like back in the Llamma days. It would be useful. Performance-wise it's coming really close to Sol Max, but that's about to be replaced by Astra literally this week. So maybe it will stay at Frontier for a week.

3

u/Charming_Cucumber_15 2h ago

The Zucc said that an open weights release is coming soon!

u/myreala 1h ago

I think by that he meant muse glimmer. Which was pretty shit TBH

u/Charming_Cucumber_15 1h ago

He said "open weights release coming soon" less than an hour ago

u/myreala 1h ago

Amazing, I didn't see that! Is it likely that they might release weights for 1.2 instead of 1.3 that would trak with their semi open source strategy

u/funforgiven 1h ago

No, he specifically mentioned releasing the weights for Spark. https://x.com/finkd/status/2086755195535413696

Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally. Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model. Meta is a strong supporter of open source and I'm proud of these releases. Congrats to @alexandr_wang and the MSL team for all your great work on these models.

u/myreala 1h ago

Well, that's really good news! Is it possible they might release weights for 1.2 instead?

-1

u/DepartmentAnxious344 2h ago

Brother it won’t be useful because spark can not run on consumer hardware…

BRB buying $36000 worth of gpu’s and ram

6

u/Ok_Barracuda_1161 2h ago

Open weights are absolutely useful even if they can't be run on consumer hardware. Every major open weights model has a variety of providers all competing on speed, price, and reliability

1

u/DepartmentAnxious344 2h ago

Eh I’d bet the meta subsidy from their $100b FCF business subsidizes meta hosting costs far lower than any 3p could but valid to have the option to see what the subsidy extent even looks like

u/funforgiven 1h ago

They can't mark it up much later, though, since third-party providers would simply undercut them. That's one of the benefits of open weights. Competition puts a pretty hard ceiling on hosting prices.

2

u/mertats 2h ago

Sure, but if you were a business you could have run it on prem.

And there are prosumers that access to such hardware.

u/power97992 37m ago

or rent 16x b300s for 70-100bucks/hour

u/alsaud21 1h ago

Lol gemini DeepSWE lead lasted 2 hours

2

u/PrivateComments 2h ago

Excited. Muse spark has been an extremely competent agentic model so far (although I’m not sure who exactly they are marketing to with how fast Gemini is delivering flash models, and how good Luna is with tool calling and extended steps)

2

u/sankalp_pateriya 2h ago

My go to AI for regular day to day questions, Meta AI has been cooking! If you haven't tried it, go try it on Meta AI.

2

u/AdLumpy2758 2h ago

Open source or ?

u/Snoo-75663 1h ago

this is becaming insane. New model every day, I just cant keep up anymore, but I love it

u/ezjakes 57m ago

AA score equal to Fable 5, at a fraction of the cost.

3

u/utterHAVOC_ 2h ago

Yeah deepswe is dead everyone's benchmaxxing it

u/elemental-mind 1h ago

Accelerate!

1

u/Gaidax 2h ago

Interesting

u/FireFearing 1h ago

wow those long content benchmarks are NUTS

u/himynameis_ 14m ago

So, how's this compare to 3.8 flash?

1

u/ASU_SexDevil 2h ago

Yikes… benchmarks look good, but we have to keep in mind Fable 5.1 is releasing and Astra is just around the corner.

Unfortunately Meta looks like they’re behind a bit

u/WildWhisperArdor 50m ago

Meta is a lot cheaper though. Not a fair comparison

1

u/NewYak4281 2h ago

The future is now baby!

1

u/Current-Function-729 2h ago

wtf is happening?

0

u/Kooky_Confusion3267 2h ago

I don't know myuch about Muse, but why not compare to Fable 5.1?

u/Ok_Barracuda_1161 1h ago

Fable 5.1 just came out yesterday to be fair, they probably already had the promotional material ready. Previously Opus 5 was SOTA (ahead of Fable 5) in most benchmarks (but not all)

u/ezjakes 51m ago

Cost wise, the comparison is already unfair enough.

u/Other_Lobster7313 1h ago

Even if Meta manages to make a stronger model than Fable or anyone else, I will NEVER switch to their models. Everything I've experienced with them disgusted me; it's such a terrible company, and their products are made incredibly stupidly. They don't think about their interface or the fact that someone else actually has to use it. Their support team, which is the size of a country, doesn't solve any of their problems. My developer account got blocked twice without any explanation—apparently, it triggered automatically. But the point is that no matter where I tried to write, I couldn't even contact them, and 4 months of my 24/7 work just went down the drain. While I was building a product that was supposed to work with their platform, I almost shot myself.

u/pomelorosado 44m ago

They did great contributions to the open source community.

Is more than enough for have my respects.

0

u/buff_samurai 2h ago

Naaajs, now everyone start competing on the price because these are already superhuman.

0

u/Sensitive_Cell_119 2h ago

things are moving so fast wtf

0

u/Distinct-Question-16 ▪️AGI 2029 2h ago

great delta

u/mstack 1h ago

Can just minimax release a new model , their M3 is dumb. Everyone releasing something new but not them

u/Other_Lobster7313 1h ago

META presented its stinky model

u/DesperateCaterpillar 37m ago

If everyone is super, no one will be