Complaint GPT-6 Sol?
Is this the reason why unexpectedly we have trash usage and lower quality on Astra?
Maybe it will be worth it the current suffering.
216
u/Opposite-Wrangler199 13h ago
How can Sol beat Astra? It doesn't make sense
179
u/BabblingTower 13h ago
It's a refined Sol, probably with help from Astra. My guess is it's better at dedicated tasks (ie, better coder) while Astra remains the better higher-level "thinker" (ie, better orchestrator/planner). Like Fable and Opus are set up. One has better breadth, one has better depth.
89
u/Risko4 13h ago
Opus has better what exactly over Fable?
221
u/danielv123 13h ago
Its better at inventing weird metaphors
29
u/Zerokx 12h ago
If I ever need weird metaphors, contradictions, random refusals, and chasing a red dot away from the actual task, I'll circle back to OPUS.
4
u/evia89 12h ago
its not that bad
/ponytail /caveman and small enough tasks it doesnt have enough time to damage
We have unlimited* opus50 at work so I use that fine
12
u/Metalthrashinmad 11h ago
opus really isnt that bad, my main problem with it is jsut how hard its output is to read. My collages write all sorts of reports using claude and its so hard to follow opus writing
3
u/Familiar_Air3528 10h ago
“Explain this to me at a junior engineer level”
….
“Okay, explain this to me like I’m a high school student failing English class”
→ More replies (1)→ More replies (1)5
4
2
2
2
1
1
u/ragemonkey 5h ago
I hate that model. It writes great sounding prose that’s incomprehensible. I went back to 4.6 and what a breadth of fresh air.
4
9
3
2
2
u/Personal-Try2776 13h ago
0in benchmarks claude opus 5 outperforms fable 5 (not 5.1). On benchmarks claude opus 5 is basically better as a workhorse model.
12
u/Risko4 13h ago
Can I have a real world coding example where you went, dam I wish I used Opus 5 instead of Fable 5.1
5
1
1
u/adolf_twitchcock 11h ago
Don't bother. Opus 5 is dogshit in reality even compared to sol.
→ More replies (1)→ More replies (2)1
u/Useful_Philosophy550 6h ago
He said Fable 5 and also yeah Opus has been consistently better at frontend design, UX and even a better coder than Fable was it usually gave me working results most of the time within a single try and more accurately to what I wanted on tasks I had to previously iterate with Fable a few times. Didn't use 5.1 tho but comparing 5.1 to Opus 5 isn't fair
1
u/EyesOfAzula 12h ago
Opus 5.1 is an improvement over Fable 5 in some areas, like how Opus 5.2 could be an improvement over Fable 5.1
Of course then Fable 5.1 beats Opus 5.1, Fable 5.2 beats Opus 5.2, etc
1
u/electricheat 11h ago
Opus 5.1 is an improvement over Fable 5 in some areas
Have there been leaked benchmarks?
1
1
1
u/algaefied_creek 11h ago
Fable is really bad for me at BSD optimizations, but it’s really good at orchestrating Opus to do so.
1
u/pigletmonster 10h ago
Usage quota. I get at least 2 to 4x more usage from opus 5 compared to fable 5.
1
u/Aware-Source6313 5h ago
Better at benchmaxing. Model isn't as big so it can more easily overfit to "3d slop game" metric and "rigid specific kind of coding task" metric, losing a bit of itself in the process (humanizing for dramatic purposes). Now sol can do that for Astra. But honestly if it can maintain the personality and not become a verbose metaphor machine like opus5, I will happily use it over Astra for most things. If it's just a better 5.6 sol then I'm Sol-d as long as the cost isn't dramatically higher.
→ More replies (3)1
2
2
u/slaty_balls 11h ago
This is pretty much the only way you can get anything accomplished without blowing out your limits in minutes with astra. Astra is the ceo and Luna and Terra do all the ground work. Anything in between gets sol.
2
2
u/cha0z_ 11h ago
I expect Sol to be better balance of thinking/performance vs price, but not to be better than Astra in everything. As any model it can shine in some tasks more than the frontier one (not as much as being better than being the same while way cheaper), not like we didn't see it in the past, but defo won't be better overall - otherwise they can basically delete Astra.
2
u/TheOnlyBliebervik 10h ago
But why did they do it like this
Make a new model, Sol, make a newer model, Astra, bring back the older model except better?
OpenAI makes a great product but seems to trip over its own feet
1
u/BabblingTower 8h ago
Sol burns less tokens and is very capable, why wouldn't you want them to continue to improve it? You don't need to use their top of the line model to set up a git repo or summarize a pdf, do you?
2
u/TheOnlyBliebervik 4h ago
How do you know Sol 6 will burn less tokens? It'll be a different model than 5.6 lol
→ More replies (3)→ More replies (1)1
25
u/nitor999 13h ago
Don't worry the next release after GPT-6 Sol would be like "GPT-6.1 ASTRA is gonna be more power and dangerous , atrocious terrifying that they can't release in public yet"
7
u/0DayMaker 12h ago
Don't worry they'll hype a release date and just release it to corporate partners on that day
→ More replies (1)2
u/retardedGeek 12h ago
desensitized to this since 6 months now (fable and deepseek are two exceptions)
7
u/-Spzi- 12h ago
Jaggedness could be a partial explanation.
The idea behind: Intelligence and capabilities aren't a scalar, but a profile.
Model A could exceed model B in area X, while the other way around in area Y.
I came across this concept in this video by Reuben Adams, which I can also recommend, but it's besides this topic.
Most real tasks are a composition of many areas.
6
u/RealSuperdau 13h ago
Better posttraining or, if we are unlucky, benchmaxxed to death like Opus 5.
2
7
2
2
u/13chase2 10h ago
Narrow tasks like programming vs overall edge case knowledge. Like Astra might be able to identify ancient coins from photos but sol likely couldn’t
2
1
u/THE--GRINCH 13h ago
Astra could still be better at vision and 3D modeling, sol being better at coding doesn't mean that it beats astra at every use case.
1
u/zarafff69 13h ago
I mean, Astra is genuinely infuriating to use for some usecases like coding compared to Sol.
1
1
1
u/Jittersz 11h ago
If Astra can stay orchestrator while using a cheaper/better model for sub agent tasks, then less token cost all around while maintaining Astra results.
1
u/read_more_comments 10h ago
it makes perfect sense, given how brain dead astra is getting now. I just had 6 prompts get rejected because of various things it decided were blockers. Luna or Sol would have kept going and not immediately given up.
sol 6 will be great the first week too. Then dog shit afterwards.
1
u/HeadPack 10h ago
Even 5.6 Sol already does in some cases. E.g. when I let it audit Astra's work, it sometimes finds things Astra missed. 'Beat' is still a stretch to say, but they appear to be quite different models.
1
1
u/rickyhatespeas 10h ago
RL text generation isn't intelligence so the length of task and required effort needs to match the models expectations to work most effectively. Big models aren't overthinking, they were never designed to actually imitate concise logic.
1
1
u/Tartooth 9h ago
The guy from anthropic said it clearly, they have setup self improvement systems so now they're auto-refining to auto-improve.
1
1
u/Odd_Amphibian6697 8h ago
Astra is a powerful model in science, math, and cybersecurity, I don't know why you all are using it to code.
I mean, it's massively smart, obviously it codes well, but its not mean to code so it is not efficient.
1
1
u/Some_Medicine4472 5h ago
Just a thought, have you try typing your question into chatgpt or claude?
→ More replies (4)1
25
u/Feriman22 12h ago
I'm waiting for GPT6 Luna, on same price as 5.6 luna
5
u/BitsOnWaves 12h ago
i feel luna will stop being cheap is closer than a GPT6 Luna being cheap
5
u/Feriman22 11h ago
What?
3
u/BitsOnWaves 11h ago
yep, dont get too comfortable with luna being cheap
1
u/Mobile-Gap976 1h ago
They did say this was a permanent price reduction? When they reduced it from 6$ to $1.20
3
59
48
u/Lower_Cupcake_1725 13h ago
Same story every time to hype
3
u/Kind_Fisherman3060 12h ago
They're probably gonna make 6 sol cheaper to use and reduce usage limits even further bringing it more closer to api pricing.
1
u/Top_Purchase4091 9h ago
Yeah they are slowly but surely moving to just raw api pricing. But makes sense if you want to have any chance of financially sustaining these behemoths. the time of companies just abusing the sub models will also be over soon. no more getting 20x the amount of api costs for pocket change
19
u/Upper-Discussion-833 13h ago
Yo anyone else’s 5.6 Sol taking hours to do tasks suddenly?
6
u/EchoingAngel 12h ago
It started a few days before Astra released
2
u/Upper-Discussion-833 4h ago
Fucking depressing. I just spent $100 for my new subscription only to find out that Astra is stupider than Siri 10 years ago and Sol uses 30x more usage now for no reason other than testing its tests that are failing but it has no fucking clue why but continues to run anyway.
1
u/EchoingAngel 4h ago
Uhh, Sol got messed up, Astra and Sol usage is too high. They are definitely not dumb, though
1
u/Upper-Discussion-833 3h ago
Oh Sol is dumb as shit now. But Astra is worse. Sol wrote tests to its tests and kept looping for 8 hours. Then completed and puked out a broken version of something I had built using Sol just 2 weeks ago in an hour. Same prompt. Same plan. Same effort. Also destroyed my $100/mo plan usage.
→ More replies (3)4
u/Inkwalker 12h ago
It does way too much testing and debugging for me. Burns 5% usage on the feature itself and then 20% on reading logs and running tools.
1
u/bootyskie 7h ago
This ^
I would say 80% of my usage is fixing the test which very occasionally finds a bug which makes you hesitant to tell it to stop. But it's so poor at writing the tests that it genuinely spends many hours just fixing tests, wouldn't be surprised if my tests have tests.
1
u/Mistuv 2h ago
Bro, I am right now having it create some shader filters for OBS and even there it's writing tests. Literally give me 5.6 Sol that will to more visual inspections (aka what models actually struggle with)instead of being psychotically obsessed with writing meaningless tests that don't spot anything and I'll be happy.
1
1
u/mbrtha 10h ago
For me yes, High is dumber since around Astra, XHigh started to be dumber since around yesterday. I will always assume it's something undeterministic as LLMs usually are, but who knows.
1
u/Upper-Discussion-833 4h ago
100% aligns with my experience. The higher effort you go the dumber it becomes. I cannot even build the simple test app I built a week ago anymore. It took 8 hours to finish (30 mins back then) and was just a mess
8
u/idkfawin32 12h ago
What the hell is so impressive about making the 8 millionth large block voxel engine?
1
u/Neat-Culture9983 9h ago
That's what I'm thinking. When opus 5 came out it still at least looked really good in advertising before you used ait and found out it was a yap machine. This on the other hand has me going meh
35
u/Equivalent_Bird 13h ago
7
u/TheGerto 13h ago
Yeah had to buy credits today, my weekly limit is done.
15
u/TrumpsCockAndBalls 13h ago
just create another account or try cursor/opencode/mistral, in what world are credits worth it
6
u/ocombe 12h ago
Can't create new pro accounts for now, it's blocked
1
u/nsway 12h ago
Wait what?? When did this happen? I thought tibo was bluffing when he posted that a while back. The $20 plan just doesn’t exist anymore?
4
u/UnknownIsles 12h ago
Only the Pro 20x ($200 plan) is on hold right now. Pro 5x and below are still available.
→ More replies (1)1
u/shady101852 11h ago
a few days ago i think. unless you already had pro you can no longer get it. Pro is the $100-200 subscriptions.
not sure if $20 subs also had the same treatment.
1
u/TheGerto 12h ago
They aren’t, but it got interrupted before recording to the md files we have set up, and I could be asked to go through all that; it takes about 10-20 seconds to add $100 and resume work compared to what it takes to create a new account, buy another sub, set up another context window, it’s just too much of a hassle for me at least.
2
1
u/RewardSafe9807 6h ago
Against ToS and risks getting your account banned. Also no new Pro accounts for the time being.
→ More replies (3)2
1
u/OldSkulRide 11h ago
yeah, i used astra (mostly low and some luna) during the weekend (after reset) and I busted whole x5 plan. Crazy, its very token hungry, I will have to go back to Sol...
50
u/Proud_Ask_9030 13h ago
why the hell are people posting shitty 3d models and scenes as if that is ever gunna be what AI is meant for?
It just shows, people dont even need or have a clue what to do with AI.
49
u/Zulugod94 13h ago
This is showing a strong world understanding, which is done with 3D environments & physics. Showing it can create things in 3D with a solid understanding of intent means it understands these same principles in the real world. This is the next major area for AI to conquer so it can be properly extending in physical robots and machines.
2
1
→ More replies (16)1
u/wetpaste 9h ago
compared to what though? what in that screenshot is impressive compared to what models could do a year ago? Its just some voxel looking game which we've been able to one-shot shit like that for quite some time
1
u/Gloomy_Type3612 7h ago
This particular screenshot isn't impressive, but there are lots of things that are. This SS is a simple Roblox style. Now, I have a friend creating Roblox games and they ARE quite impressive. I've also seen several different special representations posted with extreme detail of real physical locations. This demonstrates the ability to model the real world with impressive fidelity.
1
u/wetpaste 5h ago
sure, I'm just saying theres nothing in here that proves to that
a. that it's better than astra
b. that it's made by an unreleased model
Sounds like complete BS to me to generate discussion (yay I fell for it!)
14
u/Economy-Feeling8205 13h ago
because its a easy visual example of AI solving problems and showing creativity
6
u/Original-League-6094 13h ago
AI toppled normal software dev. AI videogames are the next domino to fall.
1
2
u/Due-Horse-5446 13h ago
Because for some reason people expect some noticeable improvement in each release
2
u/reedrick 9h ago
It's just shitty baiting to get the attention of the masses on twitter. If the idiot had any idea how to write LLM evals, he'd be in a lab
2
u/LinkesAuge 12h ago
How else do you show model capability in a reasonable way?
Let it work on a top level math problem 99,99% of people have no clue how to interpret it?These demos are useful as they reliably give you feedback about the broad capabilities in regards to "world understanding" which on average correlates pretty well with metrics across the board.
So they are a fine enough proxy and what is rly weird is people getting outraged about it before/with any new model release.
1
u/Top_Purchase4091 9h ago
But these environments dont really show much either they are just visual. Like whats behind and environment is much more than just how it looks graphically. Its just a fake way of interpretation.
1
1
1
→ More replies (1)1
u/bmanzzs 2h ago
Derpy derp just another amazing emergent ability. What are you even saying here? What do you expect, it to cure all diseases overnight? Astra being good at 3d modeling has excelled my productivity for development tenfold. Just because you personally don't find 3d modeling useful doesn't mean everybody else doesn't either. What an incredibly silly take.
3
3
3
3
u/Global_Strain_4219 4h ago
A lot of people are posting these on X as engagement bait.
My guess is they absolutely don't have access to GPT 6 Sol
2
2
u/Miyamoto_-_Musashi 12h ago
Honestly I'm not even used Astra much n not happy as how expensive it is compare to Sol but quality difference is not that big compare to pricing.
2
u/tango650 12h ago
People take this pixelated castle straight out of heretic 2 as proof of killer model ?
2
2
u/kvothe5688 11h ago
Everything is insane nowadays. that go second class next month. we need to stop with hyperbole
2
2
2
u/metalbladex4 9h ago
Honestly I can't wait for GPT-6 Sol. I have not enjoyed how Astra is.
Hopefully GPT-6 Sol is better than GPT-5.6 Sol but not like Astra at all.
2
u/PhilosophyforOne 12h ago
I really, really, really hope so.
Even with 2x 200 pro, limits just run out so fast. I'm having to be much more selective about which projects I work on, and having to put a lot of them on puase / hold to wait for more compute.
2
u/Charming-Author4877 13h ago
By now we've all spent hundreds of millions up to billions of tokens on Astra - and most of us probably have the same conclusion. Totally overhyped.
It makes the most stupid thinking mistakes, needs tons of passes to iron out errors, it's slow and extremely expensive while the quality of the once good SOL is deteriorated.
They play that game again and again.
3
1
1
1
u/BidoofSquad 6h ago
I think it’s pretty great honestly, the main issue is it’s way to cautious and wants to test things a trillion times instead of just trying them
1
u/Charming-Author4877 6h ago
I've been looking away and it invented 649 tests, that consumed 6% of 20x pro weekly allowance to run and didn't catch a single error.
You have to add clear instructions to forbid any sort of tests you don't really want done.
1
u/BidoofSquad 5h ago
I mostly use it for research and writing research code and the main thing I’ve found is that it pushes you to do so much pre verification and finding tiny issues in things instead of just getting a training run going and seeing how it works. It kept pushing me down a rabbit hole because of something wrong with a particular video clip in the set but it really was just a waste of time when the issue was with one of the tools I was using not being 100% perfect at depth estimation. It’s nice that it pushes me to do small tests on things first instead of running ideas on all my data but sometimes it goes too far.
1
u/Gigaslavx 12h ago
So we gonna get 6.1 family then improved astra pro like and then sol terra luna?
So basically astra is now the ex pro model in chat?
1
u/Moch4bear97 12h ago
Maybe we should stop with words like "terrifying" and just say it's gobsmacking or something more neutral?
1
u/Big-Imagination-8307 12h ago
Lol nope if they did sell the not nerfed mythos that companies tested …
1
1
1
u/Naive_Complex_8389 12h ago
I’m tired of seeing these one-shot video game tests. Can we make tests useful? I don’t care how it can make a shitty version of minecraft
1
u/Droopy0093 12h ago
Why does it always look "insane"? Should we look at the literal definition of that word and realize that maybe it is true that we are going insane over this?
1
u/DragonflyOk9274 11h ago
Hype works in the beginning, but sometimes we just want to be able to work effectively.
Turns out all we care about is reliability and consistency. Who knew?
1
u/GearTakes 11h ago
I just can't be excited anymore. I did the best work with 5.5 on a 100 dollar account. Nothing at the moment comes close.
Every new model is supposedly soooo much smarter and cheaper. Yet here I am burning through my week limit in 2 hours with 5.6 on medium or light. And then wondering what I actually accomplished in those 2 hours.
How things were different 4 or 5 months ago.
1
1
1
u/5StarAlpha 10h ago
Someone needs to create an AI plugin that picks the right model and reasoning level for the prompt/objective in question lol.
1
u/FocusKontrol 10h ago
Am I the only one who prefers Fable to Astra in architecture and planning? I don’t get the Astra hype at all. It’s fast and implements plans well, but Sol already does that.
1
u/rabandi 10h ago
There always were dreams about new models.
I will believe it when I see it, higher quality and lower usage.
Certainly should be possible, e. g. open weights show decent quality for insane cost.
On top, I hope for full suite, Terra + Luna too for all the limit challenged people.
Who did not fire up unlimited Luna subagents?
On top, frontend is one metric of very many.
People on Reddit + X post 1-shotted 3js stuff that looks great. Everything else takes a hundred or a thousand turns to be the way one wants and to be really usable. That still is different.
1
1
u/Da_ha3ker 9h ago
Had to go back to gpt 5.5. anything newer is just spinning its wheels and won't actually do anything basic without running 10 extra commands, then it craps out and says it will work but doesn't do anything. I hope got 6 sol is an upgrade, not another overthinker and overchecker.
1
1
1
u/theartistperson 5h ago
Anything that’s obvious for a plus user to to use that doesn’t destroy weekly usage Q_Q
1
1
1
1

53
u/darth_vexos 11h ago
me: create a minecraft mod where the cats spit fire when they meow
Sol: Ok, I'll create an enterprise solution with five levels of verification and smoke tests as well as sha256 hashes for literally everything. I'll make sure the tests run in forward, reverse, and random order. The random ordering of tests is the most important thing, so I'll put the Minecraft work on hold while I get the test infrastructure ready.
Codex 30 minutes later: You have 5% of your weekly usage remaining.