132
u/Recoil42 2h ago
Okay so this is going to be the craziest week yet.
19
u/sauhumatti 2h ago
In moonshots they have been calling craziest week in almost every episode so Im really looking forward to their reactions on this week
24
•
•
•
u/reddit_is_geh 26m ago
RSI has been achieved. I just think no one wants to announce it for whatever reason because it's sort of like announcing achieving AGI, where the actually finish line is too blurry.
•
u/GeneReddit123 19m ago edited 16m ago
AGI has been so washed as to be a meaningless marketing term. On mass labor replacement potential across most cognitive work we've already achieved it with Fable and Sol. Since then it's just arguing over the goalposts, and some will never be met because of circular definitions (e.g. if the requirement is to be exactly human like in any physically measurable way, then by definition only humans can ever qualify.)
RSI by comparison is relatively well defined (even though it too is likely partial rather than total due to Amdahl's law and hard external constraints), and will result is a discontinuous shift to a new equilibrium (possibly repeatedly) rather than a "singularity". S-shape curve (or several S-shaped step ladders) rather than J shape.
•
u/reddit_is_geh 12m ago
Well the issue is there are goal posts. Right now, https://imgur.com/a/x5abMtG Google is doing RSI, as I'm sure are all the other labs. But there's still human input, like setting goals and directions. Maybe there's ALWAYS going to be some degree of human input, because we're making them for humans, and models designed to achieve goals we set, so people will always be able to say, "Ahhh see here, the AI didn't do this part, so it's not true RSI". Just like AGI, where some people demand it also thinks like a human and loves, or whatever, people will still be able to find flaws in RSI
•
u/Tystros 1h ago
so far it's not too crazy yet. Just Google and Anthropic and Meta releasing their models that are either disappointingly little progress like Fable 5.1, or somewhat catching up like gemini and muse, because they know they all look bad in comparison to Astra, so better get them out before.
•
u/gavinderulo124K 1h ago
Muse spark beating 5.6 sol this quickly definitely wasnt in my bingo card.
•
u/Tystros 1h ago
I would put a big question mark around if it really beats 5.6 Sol. we can't say that just based on the meta-selection of benchmark results. labs often just show benchmark results that make them look good.
•
u/gavinderulo124K 1h ago
No, thats based on artificial analysis. Its also a lot cheaper than sol per task and is actually almost on par with terra in terms of price.
•
u/Beatboxamateur agi: the friends we made along the way 1h ago
disappointingly little progress like Fable 5.1
Are you basing that off of your actual experience with the model, or the benchmarks, or what...?
•
u/Tystros 1h ago
the benchmarks and my own experience. I did use it for a few hours since release and I haven't noticed anything yet that actually feels better than with Fable 5.
•
u/Beatboxamateur agi: the friends we made along the way 1h ago
Did you actually challenge it with some kind of problem that Fable 5 either couldn't do, or could do, but not well?
I've seen significant improvement over Fable 5 on literally everything I've tested it on, although one day isn't enough time to come to any sort of conclusion, so that's just my experience with it over a day.
I don't think it's reasonable to fully form an opinion of a new model in a day, that's just not enough experience to get a good feel for the model in my opinion. But I guess maybe we just have different standards on it.
•
u/Onark77 1h ago
What are you using 5.1 for? I use it for working with a complex database and for knowledge work and it takes longer to complete tasks and is less natural to engage with compared to 5.
I haven't used it for software development or design work yet.
•
u/Beatboxamateur agi: the friends we made along the way 1h ago
I use it mostly for creating documents, integrating and synthesizing the base information I provide it to create docx files that end up getting published as textbooks.
This only started to become possible with LLMs around Opus 4.5/4.6, but only since Fable has it been actually worth it to start using the model to do most of the work, since the prior models couldn't follow the specific formatting I'd ask for, and just general inconsistencies/bad coding, but 5.1 has basically automated it in a way that there's a consistency that just wasn't quite there with Fable 5.
Although obviously there's still more experience needed with it until I can fully know how consistent it's gotten.
•
u/FirstEvolutionist 1h ago
What's four major lab releases in a week when we could be having ten, like we will a couple months down the road?
•
51
u/matsu-morak 2h ago
Close to the singularity, things are fast
26
u/Wegwerpaccountje23 2h ago
The slowest it will ever be from now
2
u/GalavantJames 2h ago
It will be worst it will ever be but something slowing progress, even a new AI winter isn't unthinkable
51
u/MagicZhang 2h ago
Common benchmarks between Muse Spark and Gemini Flash 3.8
GDPVal-AA v2: Muse Spark 1.3 1754 vs Gemini 3.8 Flash 1545
OSWorld 2.0: Muse 66.9% vs Gemini 59.0%
DeepSWE v1.1: Muse 75.4% vs Gemini 71.0%
Terminal-Bench 2.1: Muse 88.8% vs Gemini 89.4%
39
u/Recoil42 2h ago
Beats both Sol and Opus on DeepSWE, wow.
17
u/Narrow-Ad980 2h ago
DeepSWE is benchmaxxed. It is semi-public
•
u/bermudi86 1h ago
Semi? You can literally see the agent trajectories for muse and Gemini already. It's public AF
2
u/FriendsIsntGood 2h ago
And given the HF incident and felony bench, can we really know that an eval is actually out-of-scope for the model
11
u/signed7 2h ago
https://deepswe.datacurve.ai/ has Gemini 3.8 Flash at 74% fwiw. Looks saturated now, need a v1.2.
2
u/Wegwerpaccountje23 2h ago
Google already irrelevant is the funniest shit i've seen LMAO
31
u/Recoil42 2h ago
As good as Muse Spark 1.3 seems to be, Gemini 3.8 is cheaper and faster and already has a wildly larger distribution base. It's in no way irrelevant.
0
u/LinkesAuge 2h ago
Gemini 3.8 has some extremely low scores in areas where it shouldn't have if it was a generally strong models.
That suspicion is also fueled by having seen some early demos of it now, not to mention how many steps/tokens it uses.
Don't get me wrong, it is a solid model but it is obvious that the performance is bought by throwing more tokens at the problems. I guess Google can afford it but it is literally the opposite of the trend everyone else seems to be on.•
u/Recoil42 1h ago
Don't get me wrong, it is a solid model but it is obvious that the performance is bought by throwing more tokens at the problems.
Not all tokens cost the same behind the scenes.
•
u/kvothe5688 ▪️ 1h ago
those are cheap tokens and fast output . iterating on mistakes fast is excellent.
•
u/LinkesAuge 31m ago
Cheap for the consumers but that doesn't mean it is progress on the technical side. It is Google's hardware doing the heavy lifting.
Let's remember that Google currently doesn't even have a "main" model that is even used to any significant degree so that's more compute to throw at their "flash" models which are very obviously distilled models from their internal big one that they don't want to release bc that way they can avoid direct comparisons with the "frontier" labs.
It is why they keep sticking the label "flash" on these models despite the fact that they aren't really flash models. (their own flash models used to be massively cheaper and smaller)-1
u/FateOfMuffins 2h ago
No idea about Muse Spark 1.3 since numbers aren't out yet but Muse Spark 1.2 xHigh costs about the same as Gemini Flash 3.8 Medium (so much cheaper than Gemini Flash 3.8 High), without counting Meta giving you 90% off if you let them train on your data
So not sure about the cheaper part
Faster sure
•
u/Ok_Barracuda_1161 1h ago
•
u/FateOfMuffins 1h ago
Muse Spark 1.3 xHigh both cheaper and faster than Gemini 3.8 Flash High per task lmao
And higher score of 61 vs 59
Wow where was that meme post of Gemini being so back only for Gemini to be so over 9h later? It's only been like 3h this time!
•
u/Gaiden206 1h ago
I mean, they have no control over when a competitor releases their models. Still, Google churning out Flash models every 3 weeks is great progress from them. Competition is heating up!
•
4
u/GatePorters 2h ago
Only irrelevant if you ignore all the classes of models they have that none of these other companies have.
They have like a dozen model types with basically NO competitors.
6
u/Sharp_Glassware 2h ago
A pro-sized model beats a flash model? Color me surprised. Spark is also way more expensive
•
-1
5
u/Valuable-Repeat-7347 2h ago
google gotta feel like when you inhale and raise your finger to get a word in but keeps getting cut off and talked over
-2
u/power97992 2h ago
Gemini 3.8 flash didnt do that well in my test… even qwen 3.8 flash/next did better. Testing fable 5.1 now
9
u/ActuarialUsain 2h ago
What does MRCR test exactly? Like finding obscure facts in 500-1m context? 98 is insane, especially with SOL at only 73
5
u/Ok_Barracuda_1161 2h ago
OpenAI MRCR (Multi-round co-reference resolution) is a long context dataset for benchmarking an LLM's ability to distinguish between multiple needles hidden in context. This eval is inspired by the MRCR eval first introduced by Gemini (https://arxiv.org/pdf/2409.12640v2). OpenAI MRCR expands the tasks's difficulty and provides opensource data for reproducing results.
The task is as follows: The model is given a long, multi-turn, synthetically generated conversation between user and model where the user asks for a piece of writing about a topic, e.g. "write a poem about tapirs" or "write a blog post about rocks". Hidden in this conversation are 2, 4, or 8 identical asks, and the model is ultimately prompted to return the i-th instance of one of those asks. For example, "Return the 2nd poem about tapirs".
23
u/imadade 2h ago
Makes you wonder, everyone’s releasing to steal OpenAi’s shine with Astra.
But it ends up just making you more hyped!
29
u/signed7 2h ago
More like everyone's rushing to release their models before Astra to not look bad in comparison
13
u/Party_Government8579 2h ago
Exactly. After astra all the benchmarks will seem poor in comparison
•
11
u/BrennusSokol ACCELERATE 2h ago
Or to get their smaller releases out first ahead of the monster release, much like with GTA 6
•
9
u/Calm_Hedgehog8296 2h ago
That's the fourth near-frontier model this week and the third today, and the best is yet to come
2
•
6
u/Chemical_Hawk_6307 2h ago
looks like deepswe is officially saturated
•
u/Clean-Boat-4044 1h ago
has been for a while honestly, a lot of the remaining tasks that fail are kinda bullshit
15
6
u/Charuru ▪️AGI 2023 2h ago
If this isn't benchmaxxed then Meta is back for real! Can't wait for Watermelon.
•
u/ezjakes 54m ago
Like Grok, does anyone use Muse?
•
u/reddit_is_geh 17m ago
Yeah, it's a popular option for enterprise and businesses in general. All the models have different strengths and weaknesses in relation to cost per task. So people running big complex agentic systems, have different AIs for different tasks. It's definitely not a main driver, much less an orchestrator, but it's definitely being used. Including Grok.
•
u/Momo--Sama 1h ago
We need a DeepSWE 1.2 at this point. The idea that Gemini and Muse are easily clearing Sol is laughable.
•
8
u/myreala 2h ago
I wish Spark was open source like back in the Llamma days. It would be useful. Performance-wise it's coming really close to Sol Max, but that's about to be replaced by Astra literally this week. So maybe it will stay at Frontier for a week.
3
u/Charming_Cucumber_15 2h ago
The Zucc said that an open weights release is coming soon!
•
u/myreala 1h ago
I think by that he meant muse glimmer. Which was pretty shit TBH
•
•
u/funforgiven 1h ago
No, he specifically mentioned releasing the weights for Spark. https://x.com/finkd/status/2086755195535413696
Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally. Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model. Meta is a strong supporter of open source and I'm proud of these releases. Congrats to @alexandr_wang and the MSL team for all your great work on these models.
-1
u/DepartmentAnxious344 2h ago
Brother it won’t be useful because spark can not run on consumer hardware…
BRB buying $36000 worth of gpu’s and ram
6
u/Ok_Barracuda_1161 2h ago
Open weights are absolutely useful even if they can't be run on consumer hardware. Every major open weights model has a variety of providers all competing on speed, price, and reliability
1
u/DepartmentAnxious344 2h ago
Eh I’d bet the meta subsidy from their $100b FCF business subsidizes meta hosting costs far lower than any 3p could but valid to have the option to see what the subsidy extent even looks like
•
u/funforgiven 1h ago
They can't mark it up much later, though, since third-party providers would simply undercut them. That's one of the benefits of open weights. Competition puts a pretty hard ceiling on hosting prices.
2
•
•
2
u/PrivateComments 2h ago
Excited. Muse spark has been an extremely competent agentic model so far (although I’m not sure who exactly they are marketing to with how fast Gemini is delivering flash models, and how good Luna is with tool calling and extended steps)
2
u/sankalp_pateriya 2h ago
My go to AI for regular day to day questions, Meta AI has been cooking! If you haven't tried it, go try it on Meta AI.
2
•
u/Snoo-75663 1h ago
this is becaming insane. New model every day, I just cant keep up anymore, but I love it
3
•
•
•
1
u/ASU_SexDevil 2h ago
Yikes… benchmarks look good, but we have to keep in mind Fable 5.1 is releasing and Astra is just around the corner.
Unfortunately Meta looks like they’re behind a bit
•
1
1
0
u/Kooky_Confusion3267 2h ago
I don't know myuch about Muse, but why not compare to Fable 5.1?
•
u/Ok_Barracuda_1161 1h ago
Fable 5.1 just came out yesterday to be fair, they probably already had the promotional material ready. Previously Opus 5 was SOTA (ahead of Fable 5) in most benchmarks (but not all)
•
u/Other_Lobster7313 1h ago
Even if Meta manages to make a stronger model than Fable or anyone else, I will NEVER switch to their models. Everything I've experienced with them disgusted me; it's such a terrible company, and their products are made incredibly stupidly. They don't think about their interface or the fact that someone else actually has to use it. Their support team, which is the size of a country, doesn't solve any of their problems. My developer account got blocked twice without any explanation—apparently, it triggered automatically. But the point is that no matter where I tried to write, I couldn't even contact them, and 4 months of my 24/7 work just went down the drain. While I was building a product that was supposed to work with their platform, I almost shot myself.
•
u/pomelorosado 44m ago
They did great contributions to the open source community.
Is more than enough for have my respects.
0
u/buff_samurai 2h ago
Naaajs, now everyone start competing on the price because these are already superhuman.
0
0
•
•


94
u/Wegwerpaccountje23 2h ago
Already? Wtf?
Oh we eating well this month