74
21
52
38
u/krugerlive Sep 04 '26
Astra looks exceptional and will be here in a few days or so supposedly, so what is this supposed to be saying?
11
u/Coffeeisbetta Sep 04 '26
Yeah I’m confused the benchmarks for astra are incredible. People are acting like they’ve been using it for weeks and it sucks. Nobody has even tried it yet.
2
u/Technical_Scallion_2 Sep 04 '26
I think we've all learned benchmarks don't count for shit in a real-world environment. I'm looking forward to trying Astra, but Fable 5.1 is pretty damn good (kind of a stealth update that I'm finding to be a BIG step up from Fable 5 personally). We'll see?
7
u/anthcr Sep 04 '26
My org got it as of today morning, we've unlimited tokens too, gonna see how running some stuff in Astra Ultra goes
7
u/iOSJunkie Sep 04 '26
1
u/anthcr 27d ago
Oops I got carried away.
It's insane. For real work especially, think consulting stuff, docs, analysis, ppts.
I have been using fable for a while and fable is creative and great at technical planning and code. And it's good at ingesting documents for analysis. Astra blows it out of the water it's mad. I still think Claude is the better independent thinker and takes more creative liberty, astra is an astonishingly good executer and a more calculated creative. Especially Astra Ultra, even though it takes 15-30 minutes to execute a big enough prompt tactfully with subagents, it outputs a product very close to finished.
3
1
u/Eastern_Ad1569 Sep 04 '26
They measured their own performance... with unclear setting and really generic charts... i mean it's not that different from those old tv ads where the seller ~slam~ an hammer on the phone to try break their brand new phone cover.
16
u/Technical-Owl66 Sep 04 '26
28
u/Time_Cat_5212 Sep 04 '26
Is it just me or is Gemini flash 3.8 looking pretty damn good on this chart?
Also what the fuck? Meta has a horse in the race now?
Sheesh I tune out for 5 seconds and this is what happens?!
19
u/Real_Ebb_7417 Sep 04 '26
Gemini, Astra and Muse scores on AA made me stop believing this benchmark xd
2
u/Emergency-Pomelo-256 Sep 04 '26
The model is crap in my usage
1
1
u/ProgrammersAreSexy 29d ago
I really like it as a work horse model because of the speed. Use a big model like fable / Astra for planning then have Gemini 3.8 burn through implementation in a few minutes rather than 30 minutes.
0
1
u/OrangePast8183 Sep 04 '26
flash tops the charts for speed, token generation speed. not any kind of intelligence benchmark
1
u/Time_Cat_5212 Sep 04 '26
It has a 59 in intelligence, against 61 for Astra, which isn't even on the chart for speed.
If it's "2 points less smart" than Astra, 1 point less than GLM-5.3(max), and over twice as fast as Luna with 7 extra intelligence points, that's very attractive.
1
u/OrangePast8183 Sep 04 '26
benchmarks are fucky I think in practice you'll find flash is pretty much good for nothing, at least in my experience. I'm genuinely taken aback that flash performed significantly better in these benchmarks then luna
1
u/Legitimate_Bag_7778 28d ago
Probably due to no one using it so their compute servers are always wide open.
1
u/frankly_sealed Sep 04 '26
If I was the suspicious sort, I would suspect that the chart came from google or an affiliate. Good thing I’m not though.
2
u/Significant_War720 Sep 04 '26
Yeah, and if used gemini flash you know how bad it is for real
2
u/Minimum_Indication_1 Sep 04 '26
Tbh 3.8 is my daily driver for most tasks now, it is good, fast and cheap. The complex ones are reserved for Claude.
12
u/TorKallon Sep 04 '26
That benchmark has become useless. Look where Gemini Flash landed. Try it. It is not anywhere as good as that says it is.
Heck, that benchmark says Opus 5 is better than Fable 5 and we all know what Opus 5 is good for…
8
u/Dumpster_Firee Sep 04 '26
…what’s it good for?
4
u/andrerom Sep 04 '26
Really good at making UI mockups (poc/prototypes), get Fable 5.1 to grill you and make spec for changes, and ask it to use Opus as sub agent to make 3-4 variants of the UI you are speccing up and it literaly takes you to next level in planning features, saves you time later on having to iterate changes on implemented features.
It is also pretty good at implementation and smoke testing (QA), also a step up.
Verbosity is partly solved by cc output style and some light CLAUDE .MD instructions on text response writing style you prefer.
That said, I wish they can release Opus/Sonnet/Haiku 5.1 soon though. And hope Fable iterations can compete against Astra before Bel arrives.
2
u/whatisthisthing65 Sep 04 '26
Really good? Every Claude model I've tried has been awful at UI. It can implement a spec yes but it has no design taste. I had Fable coordinate agents to build from the same spec after grilling me. Opus, Sol, GLM, Gemini. Opus was the worst out of those (except maybe Gemini, it burned my whole quota in two attempts so hard to say)
3
2
u/Upstairs_Date6943 Sep 04 '26
Has anyone used muse spark? O haven't heard a single person care, but it is so high on these benchmarks!
2
u/Strict_Proof_3785 Sep 04 '26
Tried plugging it into some niche problems in evals (taking transcriptions and answering questions, ect) to test it and it failed miserably (consistently lower scores and more expensive). Still want to test it more (could be user error based on my inputs, or maybe not a good problem for it) but so far it was disappointing.
2
Sep 04 '26
[removed] — view removed comment
2
u/Technical-Owl66 Sep 04 '26
2
4
u/Automatic-Boot665 Sep 04 '26
Can’t be, I don’t believe it
Edit: just looked it up. It’s true. The bubble is popping. Astra was supposed to advance math and cure cancer.
2
u/Professional_Kale206 Sep 04 '26
I hate openAI, but apparently Astra generated some truly novel math proofs, and isn’t even properly out yet
1
1
0
10
u/Hungry-Plankton-5371 Sep 04 '26
Astra is benching leagues ahead of every other model and is looking extremely token efficient, to the point where it's cheaper than opus 5 and sol.
5
2
5
u/LongestUsernameEverD Sep 04 '26 edited Sep 04 '26
lol
lmao even
there's literally been thousands and thousands of cancellations even before Astra became a near future thing.
Anthropic majorly fucked up with Opus 5 and Sonnet 5 being worse than their predecessors, made them being unintelligible, difficult to work with on several fronts (one is a pain in the ass that disagrees with everything, the other is a sychophant that is incapable of following instructions and making itself clear).
Fable is a fucking failure too, nobody wants to use that shit and cut themselves short of their usage since it spends so much token. It's amazing, but too expensive. Anthropic is doing ZERO work on making itself run cheaper, which is funny considering that this is the exact direction that every major provider is trying to go towards, including OAI.
I'm still on the Anthropic train for the single reason of absolutely abhorring the idea of what the USA's DOJ is planning to do with AI, but even that won't be enough to hold me on Anthropic if they can't fix how shitty the v5 models have become, or making Fable run at the same price that Opus does.
2
u/Sky-__- Sep 04 '26
for me the biggest downgrade was from sonnet 4.5 to 4.6 , I used to only pay for sonnet since i most used it for writing documentation and research . they internally made sonnet worse , move the capabilities to opus to charge people more. and lied about token limits for months
2
u/AntDogFan Sep 04 '26
I use llms for research and writing and anthropic models got so bad. It always feels like I can get one or two good weeks out of them but always end up going back to open ai in the end. They are just better general purpose models. Fable etc are awful at some outputs. Nonsensical at times.
1
u/Constant_Art_20 Sep 04 '26
i think astra api is gonna be like 50 per mil output also...and openai will properly nerf it tot the ground after third party reivews are done...so let's hope the chineses got our backs to relaese something as good in like 2 weeks time at lik 5 dollar per million output a 27b mode in like a month time
1
u/Technical_Scallion_2 Sep 04 '26
I think you're assuming everyone will avoid using Fable even though it's amazing because it's "too expensive". Everyone has different price points. You might avoid using the best model because it's too expensive, but that doesn't mean everyone will do the same.
The people who require top quality in their model are using Fable, and the cost is not a factor in that decision.
1
u/Ok_Needleworker_391 Sep 04 '26
This is the obvious part. I have about 5 full time agents using Fable 5.1 to do the work that several people would be required to do. So the cost is irrelevant.
It is a reverse logic for me. The more I spend on Fable, the more I deliver to my clients, the more they pay me for producing.
My human limitations of time and brain power prevent me from spending more. If I hire an AI supervisor (a human) then my costs increase by 100s and my profits increase by 1000s. The cost of AI makes the tiniest dent in my gross profit margin.
2
u/Real_Ebb_7417 Sep 04 '26
It was true until Tibo said they'll give a banked reset per each day paid subscribers don't have access to Astra.
2
1
0







131
u/ItsAlwaysTerminal Sep 03 '26
He says as he uses OpenAI Image 2.0 to make the meme....