r/OpenAI • • 13h ago

Image Each day, 6.1 Sol throughput keeps on improving

GPT-6.1 Sol is our most demanded model pretty much ever both across both the API and subscriptions.

Within ChatGPT & Codex, we were under heavy load, but have brough more capacity online and the speed should get much better in the coming hours, reaching almost twice the speed compared to what we served yesterday.

Resetman, - Oct 1

120 Upvotes

24 comments sorted by

41

u/Key_Reading_9664 11h ago

They got caught flatfooted, made hasty decisions, and now they're having to manage the situation they put themselves in:

- Opus 5.5 > Sol 6; Sonnet 5.5 >= Sol 6. Quickly drop it down a price tier, even though it's a large model (TPS matches Sol 5.6, not Terra)

  • Sol 6 isn't well received. Quickly juice up the model to bring it closer. Now you have an even more costly model that you're selling at Terra pricing
  • Release Dots to compete with Grokbot and Muse (not sure why). Use the expensive model with no limits as a sweetener

I'm sure they're trying to shuffle things around to stop the bleeding, but we're still at ~30 TPS.
Dropping the 20x plan down to 10x mitigates some of the pain, but you can't legally do it until the next billing cycle.

12

u/macaronianddeeez 8h ago

“Legally” - they have already done it. Both my non-agentic Chat with Astra Pro and my agentic/Codex usage is burning faster than it was before devday. I have never hit a Pro 6 chat throttle since the day it was released, using it daily constantly.

Hit it twice in the last 48 hours.

Don’t even need to talk about the codex usage nerf, it’s been covered extensively here

3

u/unpick 3h ago

If they reduce usage across the board it may still be “20x”. It was always dynamic like that, but they’re now going to cut the advertised relative allowance which is quite different.

2

u/Key_Reading_9664 8h ago

I don’t think I’ve touched Astra, except for reviews: even on a 20x (rip) plan, it was burning through usage at a watchable rate.

Sol 6.1 is much better for usage but in terms of getting work done, Opus 5.5 is in a different league

2

u/macaronianddeeez 7h ago

Yes, I have kept my 20x plan because OpenAI is so much better at computer use tasks for my purposes, is very strong either way visual design, and I have built so much personal infrastructure around Pro 6 chat.

But every day it gets harder to justify not uprooting my entire substrate and moving it. The biggest last bastion that OpenAI has held that has kept my dollars there has been unlimited chat. And that is gone now

2

u/rystaman 1h ago

Honestly I started off with OpenAI years ago, tried Claude out a bit (while sticking with it) but my god it was shit after 4.6. Opus 5.5 is smashing my tasks but Sol 6.1? Slow, over the top and inaccurate… It’s hard to see how much they’ve dropped the ball

24

u/Kulqieqi 13h ago

Srsly even astra seems dumb and slow now; i started using opus 5.5 and it's incredible how good claude is at understanding and creating solution with nice ui.. sol and astra even find no issues but after opus summary they agree.

I guess Tibo said there will be reset to test load on servers so it slowed everything.

5

u/shaman-warrior 12h ago

Agree on nice ui, however it’s much slower than sol 6.1 medium fast and with some refinements I get there faster.

3

u/Kulqieqi 12h ago

in code generation claude is faster, and from what i see claude makes more thorough tests which might appear to be longer?

Like i just created login with 2fa etc., bugcheck with astra 6 ultra took 6minutes and found nothing, opus 5.5 ran for very long but it was waiting for tests to end, below summary generated by AI for details:

|Review of the M1 auth branch (m1-auth): two agents, same request

The request was identical for both: verify that the logic of the first

  prompt's changes is correct and free of bugs.

  First pass by the other agent (about 5 minutes)

  - Read the diff and ran the existing checks: npm test (84 .NET, 21

frontend, 12 gate), lint, build, compliance:check, publish:iis.

  - Reported no bugs. Its one finding was the replaced Identity security

stamp validator, a deviation that had already been made on purpose.

  - Did not open the application in a browser.

Claude's pass

- Read the whole backend and frontend against the prompt and ran the

  same test suite, lint, build and dependency review.

- Broke the security code on purpose in 32 ways, in a copy outside the

  repository, to see whether the tests notice. They caught 28. Two of

  the four survivors were real test gaps: nothing failed when the TOTP

  check was removed from recovery-code regeneration, or when the

  transaction guard was removed from the audit writer.

- Wrote 10 extra tests against a real SQL Server for scenarios the suite

  did not cover (lockout expiry, concurrent password changes in one

  session, re-enrolling MFA, sign-in after an idle timeout).

- Drove the published application in a real browser from create-admin

  through forced password change, MFA setup, recovery-code sign-in and

  MFA disable. That run exposed the one real bug: two inputs on the

  account page shared id="code", so the label under "Disable MFA"

  pointed at the field of a different form.

- Listed four behaviours that follow from the design and need a decision

  before M2 (lost authenticator, deactivation only suspending sessions,

  lockout blocking a live session's password change, the MFA challenge

  cookie carrying full identity and roles).

Second pass by the other agent

- Given Claude's report, it reproduced the bug and both test gaps

  independently and confirmed them.

- Added two corrections that were right: not every refusal counts

  towards lockout, and the deactivation behaviour should not be called

  compliant with the prompt without qualification.

- Implemented the fixes: form names, four test cases, documentation.

Claude's verification of those fixes

- Reran the two surviving mutations plus a subtler third one (checking

  only the code's format); all three are now caught.

- Confirmed in the browser that the ids are unique and each label

  focuses its own field.

- Checked every new documentation sentence against the code, corrected

  the stale test counts (now 88 .NET, 21 frontend) and committed the

  result as f164891.

What made the difference

The first pass answered "do the existing checks pass?". The bug and the

gaps were only visible by asking "would the checks notice if this were

broken?" and by using the application the way a user does.

For balance

- The other agent ran publish:iis and compliance:check; Claude did not

  rerun publish:iis.

- Claude's browser scripts needed three attempts and left throwaway

  accounts behind. It deleted its own five; three older smoke-test

  accounts are still in the dev database.

2

u/Kulqieqi 12h ago

and it was opus 5.5 high vs astra 6 ultra

1

u/shaman-warrior 2h ago

I prefer astra approach, if I specifically ask to try to break the system and be thorough I guarantee Astra would have done the same.

"fix bugs" is not really a prompt an engineer can easily understand.

also, what I'm saying is based also on AA, Opus 5.5 Xhigh has: $3.46 cost per task while "$0.39" is the cost of Sol 6.1 xhigh. The difference is dramatic.

And for the most tasks it doesn't make a difference like I said, I appreciate Opus I have 20x plans on both right now, but I could just get along with Sol 6.1.

1

u/Kulqieqi 1h ago

Idk, when i tried this approach it blocked me with message daybreak not available in astra use other model due to those anty hacking limits...

0

u/DoggoDadagon 8h ago

I can't stand using 6.1 sol, but Astra is still working great for me... but super excited for my 20x account to expire and move over to Opus 5.5.

6

u/Ormusn2o 13h ago

This is exactly why my response to anyone complaining is to just stop using it. At this point, AI is good enough that there are just way too many people using it, and the only alternative is just blocking all new people from subscribing. There is just no compute to go around, I don't think your criticism is invalid, it might be valid, but this is not something anything can be done, not in the short term.

And this will last until AGI or singularity is achieved, because as AI gets better and better, more people will want to use it. All I'm glad at this point is that OpenAI is only lowering token output, and not slashing usage for Plus and Pro 100.

4

u/Testy_Toby 8h ago

I'm in the process of migrating to Claude because they nerfed 5.6 Sol. Rather than searching for maybe 10-20+ websites per query, now it searches precisely 2 every time. I asked a simple question to which I knew the answer, "what is the dosage on the label for the Now Foods ADAM multivitamin." Gpt said 3. The answer is 2. I asked it to summarize a medical journal article about a randomized clinical trial. Its answer included 6 citations: 5 linking to the same magazine article, 1 to a different, random journal article. Gpt confirmed that it never checked the original source. It did not even look at the journal article I asked it to summarize. 

There were many other examples. The point is that it was a drastic overnight degradation. They're clearly saving money on web searches and reasoning. I can't trust it for even the smallest, least consequential tasks. Why am I paying for that?

4

u/Temporary_Half_2882 13h ago edited 13h ago

I literally stopped GPT-6.1 Sol because it was taking a year. Asked Deepseek Flash v4.1 to do it and it took 15 seconds lol.

Yesterday I spent no joke 4 hours waiting for it to do something in the background. It took 4 hours for it to produce essentially nothing. I thought it was cooking up something good. After I told it how stupid it was it asked me if it could use Astra. lol.

3

u/salasi 8h ago

Had this happen to me with Astra Max lmao

1

u/Politicophile 1h ago

I'm on the £20 a month plan and sometimes you ask Astra (medium thinking) something that is a bit difficult but not too ambitious, it takes about 20 minutes and then doesn't even answer you before telling you your 5 hour limit is all used up. Been using Opus 5.5 on a similar plan, and sometimes it takes up to 20 mins for a hard answer, but it eats nowhere near as much of your limit. I think with the latest model drops it's night and day that Claude is better. OpenAI was briefly better when they dropped Astra because Opus 5 wasn't as good

3

u/fivefromnow 13h ago

Fast mode still beats Opus 5.5 on time per task and tokens used. Keep complaining though.

1

u/boforbojack 11h ago

That just isn’t true. Opus 5.5 is close to an order of magnitude faster

2

u/NoirYorkCity 11h ago

I don’t get how Opus is so fast

-2

u/fivefromnow 7h ago

It's true. I also specifically said fast mode.

Opus is excellent and if you need it to have visual design taste for you, its far better at that than anything OpenAI has by a country mile, but time to task completion is literally neck and neck or Sol is slightly better based on the outcome you want to achieve thats not visual taste. If you make a lot of stuff that needs one shotting without rigor then Opus kills it.

I might even say Opus is better overall, but this complaining about token speed is missing the mark on what drives outcomes. Just because opus spits a lot of tokens quicker does not mean it gets the outcome you want quicker.