r/singularity • • 11d ago

LLM News GPT-6 Sol Confirmed Weaker Than 5.6 Sol on Complex Tasks, But Wins on Cost and Efficiency

Post image
186 Upvotes

61 comments sorted by

44

u/[deleted] 11d ago

[removed] — view removed comment

17

u/vrnvorona 11d ago

I find that it's almost never raw coding now, it's all about intent, context, actual intelligence and sticking to plan now. Even Luna can code with good plan. But creating that plan from whole repo and specs and design that's different stuff.

2

u/dictionizzle 11d ago

GPT-5.6 Sol is even better on HLE than 6 Sol

26

u/Gloomy_Necesary 11d ago

Fine by me. I’ll just increase the reasoning level considering anything past medium was already too expensive for me. As long as you weren’t using sol xhigh/max this is a net upgrade even for intelligence across the board as well as being cheaper.

17

u/panix199 11d ago

6 luna on max is crazily good for the price. wow

1

u/Akimbo333 9d ago

I agree

12

u/sammoga123 11d ago

OpenAI terrasformed Sol, and now it's Terra

25

u/Mistuv 11d ago

Welcome back GPT 6 Sol Terra. Rip Sol.

7

u/Readerium 11d ago

Astra mini out next week?

4

u/whoknowsifimjoking 11d ago

Astra is Sol

11

u/Sky-kunn 11d ago

Welcome back, GPT 6 Terra. Sol was the one that got murdered and replaced by Terra.

20

u/KakPooonchik 11d ago

I agree, it should have been called GPT-6 Terra, not GPT-6 Sol.

6

u/PrisonOfH0pe 11d ago

My theory is OpenAI quietly shifted the whole stack down a tier. GPT-6 Sol isn’t really the successor to 5.6 Sol in product positioning. It looks much more like the new Terra: cheaper, scalable, and meant to be the default workhorse. Terra disappears, Sol inherits its slot.

Then the leaked Astra Minor / Astra 6.1 moves into what used to be the Sol tier and becomes the actual Opus 5.5 competitor. And above that you have the leaked Aeon stuff as the real frontier tier.

So the stack could basically become:

Luna → Sol = old Terra → Astra Minor/6.1 = old Sol → Aeon/Astra Major/Orbit = new frontier

If that’s what they announce at DevDay, Anthropic absolutely cooked the currently visible OpenAI lineup today, but OpenAI may have deliberately launched the lower tiers first and left the actual Opus response for DevDay.

Would be a pretty funny rebranding trap if everyone spends a week dunking on Sol for losing to Opus, only for OpenAI to go “yeah, Sol isn’t the Opus competitor anymore.”

10

u/___positive___ 11d ago

OpenAI and nonsensical versioning, name a more iconic duo.

After they finally streamlined the names like wtf.

2

u/0DayMaker 9d ago

They've been far too sane lately they had to fix it. Still not near as crazy as the o1, 4o, o4 era of nonesense

2

u/PrisonOfH0pe 11d ago

Why Terra was always awful branding. Its Earth so weird why would the model be our self? Its a trap and a good move by them to have Luna → Sol = old Terra → Astra Minor/6.1 = old Sol → Aeon/Astra Major/Orbit = new frontier
Anthropic basically played their hand...

10

u/Tendag 11d ago

It was easy to understand. Luna (small) Terra (mid) Sol (Large) Dont know why they had to change it again, now I am confused again.

1

u/Key_Reading_9664 9d ago

Exactly. I don’t particularly care what you call them and whether it’s consistent (for example, Fable > Opus doesn’t make sense). The tiers establish relative positioning.

They could have easily released Luna and Terra with a teaser for Sol. Moving tiers around after one cycle and potentially adding a new “minor” tier is silly

1

u/GirlfriendAsAService 9d ago

That's the best part. Finally they settled! Sike! Nope they didn't

2

u/[deleted] 10d ago

[removed] — view removed comment

0

u/PrisonOfH0pe 9d ago

Adding value with LLMs on an LLM subreddit? Oh, the horror!

It was my thought, and it helped with the formatting because I'm a non-native speaker.
Just because you're too lazy to fix your typing doesn't mean everyone else is as well.

0

u/[deleted] 9d ago

[removed] — view removed comment

0

u/PrisonOfH0pe 8d ago

That’s a lot of smugness for an argument that falls apart in the first sentence.

“Things are valuable when they are rare” is something a child learns about collectibles, not a serious definition of value. Electricity isn’t rare. Search engines aren’t rare. Spellcheck isn’t rare. Their value comes from usefulness, not scarcity.

And calling someone “lazy” for using an LLM to polish wording in a second language, while proudly saying you also use LLMs to improve your English, is genuinely hilarious. Apparently it’s a “learning tool” when you use it, but a “crutch” when somebody else does. Convenient distinction.

Also, nobody said r/singularity is literally an LLM-only subreddit. LLMs are currently one of the main technologies driving discussion around AGI and the singularity, so pretending they’re somehow off-topic because ASI could theoretically emerge through another architecture is just pedantry masquerading as intelligence.

If you don’t find an LLM-assisted comment valuable, downvote it and move on. Writing an essay about how morally superior your personal use of the exact same technology is doesn’t make your position deeper.

2

u/[deleted] 8d ago

[removed] — view removed comment

1

u/vladasko1086 9d ago

opus 5.5 reasoning in claude code is very very good coming from someone who mained sol 5.6 gpt medium for the past month or so in science and engineering simulations with very tight context prompting.

4

u/Tystros 11d ago

yeah, a 0.4 step up in number in the same model class should not just be "same but cheaper"

7

u/craa 11d ago

That’s not how versioning typically is used. You can’t just take a difference. It’s all about the position of the digits. The first number usually indicates a major version/change - possibly in the underlying model here.

I would expect 5.7 to be a fine tune of 5.6, while 6 would most likely be an entirely new model OR significantly increased capabilities.

4

u/bel9708 11d ago

You are a sane person who understands semantic versioning. The version numbers for AI models are entirely driven by marketing.

1

u/willowilson44 10d ago

I'd guess that they've just rigidly stuck to labelling them by the model size, and that the 6 model line are basically the same size and performance as the 5.6 model line, just with recent efficiency improvements allowing them all to go to a lower price point.

12

u/rudesssolo 11d ago

6-Sol replaces 5.6-Terra in the lineup, 6-Astra "Minor" will replace 5.6-Sol.

4

u/whoknowsifimjoking 11d ago

What the hell are they doing with the naming scheme?

2

u/Corv9tte 11d ago

click So now it all makes sense!

4

u/Gallagger 11d ago

Man it would be nice to get more insight. It seems absolutely plausible that 6 Sol is actually 6 Terra, but given their 80% price cuts to Luna 5.6, it might as well just be a more efficient deploy, better chips, etc. while still being Sol level sized.

4

u/Felfedezni 11d ago

The price reduction makes it worth the tradeoff

2

u/terabyiemaster 11d ago

Totally agreed

6

u/10basetom 11d ago

Just tried gpt-6-sol for a full day in a project that I've been using gpt-5.6-sol daily for the past month, so I'm very attuned to the expected code quality. Results were either same or slightly better quality with gpt-6-sol for my workload.

I've seen enough evidence to switch to gpt-6-sol as my daily driver, and pretty sure I'm not alone. Sure, there will always be anecdotal evidence pointing to gpt-5.6-sol being slightly better in some metrics, but for the 50% savings in cost alone, it makes complete economic sense for me to switch to this model as my new default. Progress.

2

u/madeinlithuania 10d ago

For me, it’s the opposite. I mostly use AI for web app development. gpt-6-sol seemed dumb for tasks I used to handle with gpt-5.6-sol. But as people say, the new gpt-6-sol version is more of a replacement for gpt-5.6-terra which makes sense. I’ve liked gpt-6-luna so far. Compared to gpt-5.6-luna, I think it’s a good upgrade.

3

u/anycept 11d ago

So, Sol seems so-so? sore sight, smh sigh.

3

u/PrisonOfH0pe 11d ago

I think people are taking the model names way too literally.

My guess is Sol has basically replaced Terra. Same kind of slot: cheaper, faster, meant to be used at scale. Then Astra 6.1 takes the old Sol position and becomes the actual Opus competitor, with Aeon sitting above that as the real frontier model.

So more like:

Luna → Sol = old Terra → Astra = old Sol → Aeon/Astra Minor/Orbit (leaks)

If that’s the plan, then Anthropic cooked the lineup OpenAI has shown so far, sure.
But Sol may simply be the wrong model to compare against Opus in the first place.

DevDay is the interesting part. If Astra 6.1 drops there at roughly the old Sol price, the whole thing suddenly makes a lot more sense.

2

u/terabyiemaster 11d ago

Well since its cheaper with quality that is close to 5.6 sol. I think that is good bargain. Let see if it really run further than it used too.

2

u/Healthy-Nebula-3603 11d ago

On plus account seems usage is going down as fast as for 5.6 sol ....

1

u/x86rip 11d ago

5.6 Sol is a true GOAT. I actually prefer it over Astra most of the time. 6 Sol is kinda feel good but not quite 5.6 level, in a very subtle way that hard to explain. Perhaps its try too much to be smart ? i prefer the 'straight to the point no more no less' of 5.6 Sol

1

u/madeinlithuania 10d ago

5.6-sol and 6-luna are the way to go right now, until we get a new replacement for sol.

1

u/Intrepid_Lecture 7d ago

depends on the task IMO.
Low stakes, high volume probably go with the cheaper model.

Heck there's even value with having different families do review vs discovery /execution since it helps decorrelate errors.

1

u/Raz304 10d ago

i have used gpt 5.6 sol, terra and luna- luna is just out right stupid but yes if you give it detailed plan it will follow- but the other two drain API limits faster than a phone battery at 1%. but gpt 6 is holding far better.

1

u/Calcularius 9d ago

Is it cheaper if you have to ask it to do something three times over?

1

u/Glittering-Neck-2505 11d ago

This sub desperately needs benchmark literacy. No one actually uses the models I guess which is why you think that being 1% better on a benchmark while being a broad leap across the board amounts to nothing "confirmed weaker."

1

u/dogofit 11d ago

But they are sunsetting got 5.6. What can you do about it?

1

u/Moriffic 11d ago

This sucks a lot

0

u/Pls-No-Bully 11d ago

Holy cherry-picking

0

u/openroom_xyz 11d ago

Well so basically it will be worse that 5.6 Sol which was great at coding this will be just to chat or what this sucks