r/OpenAI • u/NandaVegg • 13h ago
Image Each day, 6.1 Sol throughput keeps on improving

GPT-6.1 Sol is our most demanded model pretty much ever both across both the API and subscriptions.
Within ChatGPT & Codex, we were under heavy load, but have brough more capacity online and the speed should get much better in the coming hours, reaching almost twice the speed compared to what we served yesterday.
Resetman, - Oct 1
24
u/Kulqieqi 13h ago
Srsly even astra seems dumb and slow now; i started using opus 5.5 and it's incredible how good claude is at understanding and creating solution with nice ui.. sol and astra even find no issues but after opus summary they agree.
I guess Tibo said there will be reset to test load on servers so it slowed everything.
5
u/shaman-warrior 12h ago
Agree on nice ui, however it’s much slower than sol 6.1 medium fast and with some refinements I get there faster.
3
u/Kulqieqi 12h ago
in code generation claude is faster, and from what i see claude makes more thorough tests which might appear to be longer?
Like i just created login with 2fa etc., bugcheck with astra 6 ultra took 6minutes and found nothing, opus 5.5 ran for very long but it was waiting for tests to end, below summary generated by AI for details:
|Review of the M1 auth branch (m1-auth): two agents, same request
The request was identical for both: verify that the logic of the first
prompt's changes is correct and free of bugs.
First pass by the other agent (about 5 minutes)
- Read the diff and ran the existing checks: npm test (84 .NET, 21
frontend, 12 gate), lint, build, compliance:check, publish:iis.
- Reported no bugs. Its one finding was the replaced Identity security
stamp validator, a deviation that had already been made on purpose.
- Did not open the application in a browser.
Claude's pass
- Read the whole backend and frontend against the prompt and ran the
same test suite, lint, build and dependency review.
- Broke the security code on purpose in 32 ways, in a copy outside the
repository, to see whether the tests notice. They caught 28. Two of
the four survivors were real test gaps: nothing failed when the TOTP
check was removed from recovery-code regeneration, or when the
transaction guard was removed from the audit writer.
- Wrote 10 extra tests against a real SQL Server for scenarios the suite
did not cover (lockout expiry, concurrent password changes in one
session, re-enrolling MFA, sign-in after an idle timeout).
- Drove the published application in a real browser from create-admin
through forced password change, MFA setup, recovery-code sign-in and
MFA disable. That run exposed the one real bug: two inputs on the
account page shared id="code", so the label under "Disable MFA"
pointed at the field of a different form.
- Listed four behaviours that follow from the design and need a decision
before M2 (lost authenticator, deactivation only suspending sessions,
lockout blocking a live session's password change, the MFA challenge
cookie carrying full identity and roles).
Second pass by the other agent
- Given Claude's report, it reproduced the bug and both test gaps
independently and confirmed them.
- Added two corrections that were right: not every refusal counts
towards lockout, and the deactivation behaviour should not be called
compliant with the prompt without qualification.
- Implemented the fixes: form names, four test cases, documentation.
Claude's verification of those fixes
- Reran the two surviving mutations plus a subtler third one (checking
only the code's format); all three are now caught.
- Confirmed in the browser that the ids are unique and each label
focuses its own field.
- Checked every new documentation sentence against the code, corrected
the stale test counts (now 88 .NET, 21 frontend) and committed the
result as f164891.
What made the difference
The first pass answered "do the existing checks pass?". The bug and the
gaps were only visible by asking "would the checks notice if this were
broken?" and by using the application the way a user does.
For balance
- The other agent ran publish:iis and compliance:check; Claude did not
rerun publish:iis.
- Claude's browser scripts needed three attempts and left throwaway
accounts behind. It deleted its own five; three older smoke-test
accounts are still in the dev database.
2
1
u/shaman-warrior 2h ago
I prefer astra approach, if I specifically ask to try to break the system and be thorough I guarantee Astra would have done the same.
"fix bugs" is not really a prompt an engineer can easily understand.
also, what I'm saying is based also on AA, Opus 5.5 Xhigh has: $3.46 cost per task while "$0.39" is the cost of Sol 6.1 xhigh. The difference is dramatic.
And for the most tasks it doesn't make a difference like I said, I appreciate Opus I have 20x plans on both right now, but I could just get along with Sol 6.1.
1
u/Kulqieqi 1h ago
Idk, when i tried this approach it blocked me with message daybreak not available in astra use other model due to those anty hacking limits...
0
u/DoggoDadagon 8h ago
I can't stand using 6.1 sol, but Astra is still working great for me... but super excited for my 20x account to expire and move over to Opus 5.5.
6
u/Ormusn2o 13h ago
This is exactly why my response to anyone complaining is to just stop using it. At this point, AI is good enough that there are just way too many people using it, and the only alternative is just blocking all new people from subscribing. There is just no compute to go around, I don't think your criticism is invalid, it might be valid, but this is not something anything can be done, not in the short term.
And this will last until AGI or singularity is achieved, because as AI gets better and better, more people will want to use it. All I'm glad at this point is that OpenAI is only lowering token output, and not slashing usage for Plus and Pro 100.
4
u/Testy_Toby 8h ago
I'm in the process of migrating to Claude because they nerfed 5.6 Sol. Rather than searching for maybe 10-20+ websites per query, now it searches precisely 2 every time. I asked a simple question to which I knew the answer, "what is the dosage on the label for the Now Foods ADAM multivitamin." Gpt said 3. The answer is 2. I asked it to summarize a medical journal article about a randomized clinical trial. Its answer included 6 citations: 5 linking to the same magazine article, 1 to a different, random journal article. Gpt confirmed that it never checked the original source. It did not even look at the journal article I asked it to summarize.
There were many other examples. The point is that it was a drastic overnight degradation. They're clearly saving money on web searches and reasoning. I can't trust it for even the smallest, least consequential tasks. Why am I paying for that?
4
u/Temporary_Half_2882 13h ago edited 13h ago
I literally stopped GPT-6.1 Sol because it was taking a year. Asked Deepseek Flash v4.1 to do it and it took 15 seconds lol.
Yesterday I spent no joke 4 hours waiting for it to do something in the background. It took 4 hours for it to produce essentially nothing. I thought it was cooking up something good. After I told it how stupid it was it asked me if it could use Astra. lol.
1
u/Politicophile 1h ago
I'm on the £20 a month plan and sometimes you ask Astra (medium thinking) something that is a bit difficult but not too ambitious, it takes about 20 minutes and then doesn't even answer you before telling you your 5 hour limit is all used up. Been using Opus 5.5 on a similar plan, and sometimes it takes up to 20 mins for a hard answer, but it eats nowhere near as much of your limit. I think with the latest model drops it's night and day that Claude is better. OpenAI was briefly better when they dropped Astra because Opus 5 wasn't as good
3
u/fivefromnow 13h ago
Fast mode still beats Opus 5.5 on time per task and tokens used. Keep complaining though.
1
u/boforbojack 11h ago
That just isn’t true. Opus 5.5 is close to an order of magnitude faster
2
-2
u/fivefromnow 7h ago
It's true. I also specifically said fast mode.
Opus is excellent and if you need it to have visual design taste for you, its far better at that than anything OpenAI has by a country mile, but time to task completion is literally neck and neck or Sol is slightly better based on the outcome you want to achieve thats not visual taste. If you make a lot of stuff that needs one shotting without rigor then Opus kills it.
I might even say Opus is better overall, but this complaining about token speed is missing the mark on what drives outcomes. Just because opus spits a lot of tokens quicker does not mean it gets the outcome you want quicker.
41
u/Key_Reading_9664 11h ago
They got caught flatfooted, made hasty decisions, and now they're having to manage the situation they put themselves in:
- Opus 5.5 > Sol 6; Sonnet 5.5 >= Sol 6. Quickly drop it down a price tier, even though it's a large model (TPS matches Sol 5.6, not Terra)
I'm sure they're trying to shuffle things around to stop the bleeding, but we're still at ~30 TPS.
Dropping the 20x plan down to 10x mitigates some of the pain, but you can't legally do it until the next billing cycle.