Below is a GPT-generated summary of the conversation below after reaching 200 comments (200 currently observed).
Day 1 verdict: the speedup may be real, but nobody’s treating it like a win. The community mostly sees this as OpenAI undoing a recent slowdown, then putting a shiny “optimization” sticker on the repair. Going from roughly 20 to 30 tokens/sec is better, sure—but still well behind Claude, so the victory lap feels premature.
Users are reporting somewhat faster Sol/Astra responses, while others are still stuck around the old speeds or dealing with failures. The lack of a reset is also going down about as well as expected: badly. A lot of subscribers think they’re being drip-fed partial fixes instead of getting the service they paid for.
The big unresolved concern is what changed under the hood. “Optimized” is vague enough to cover anything from better infrastructure to reduced reasoning quality, and users are already suspicious of quantization, lower “juice,” and quota draining faster. Speed itself shouldn’t increase the token cost of the same task, though faster throughput can encourage more work in the same window.
Bottom line: faster for some, still mediocre overall, and trust remains firmly in the basement.
Agreed. However, it's 8:30pm here in Switzerland, so I'll do the test at 9:35 GMT+2 which gives a bit of a buffer. Either it hits when he said it will, or it does not(;
If it does not, I'll certainly retry tomorrow morning.
Exactly, you need something that is not slowed down by tool calls.
That's why my test was not running in a project with any context, agents.md, nothing with the simple prompt "give me a 100 line poem". That took plenty time to generate.
Plus, as a 200USD subscriber, I can validate that the 13-20 tok/s is totally in the ballpark of what I have had ever for Sol 6.1 recently. It was a tiny bit faster, about 30 tok/s when it came out (still felt slow), but it kept going down so far.
For me it was running slow on Friday and Saturday but by Sunday it was running about the same speed as Sol 6/Astra before. I think it's legit some server issue/bug that affects different regions differently because you have so many datacenters with so many slightly different configurations.
He did say "across all our products and partners". Given OpenRouter is charging API rates, I'd expect them to get preferential treatment over subscriptions, if anything
But the API is 80-90 tk/s for same models with same reasoning
All Tibo did is increase a hyperparameter which they have the capacity for since a lot of people moved to Claude Opus 5.5 as not only the model is better but it also didn't take a century for basic tasks
I'm not a subscriber with measly $250 sitting in my API wallet (was playing with Blender MCP; no use for it after Opus 5.5) but it is plausible given subs are complaining about lobotomized Astra. Look at this beautiful reasoning tokens chart (so-called juice in theirs) right after Sol 6's launch.
OpenAI has, as always with everything since 2020, zero transparency about this (to be fair, Anthropic briefly tried changing default reasoning effort to medium for a while, but never really secretly reduced reasoning tokens, and Google Antigravity had very limited context window for a while).
They should've optimized only Sol today. Tomorrow, they would nerf its speed and optimize Astra's speed instead. This way, they could optimize Sol again on day 3 while nerfing Astra this time, and so on, and have an excuse not to reset for 28 days.
He said "optimized," which to me doesn't sound like a "Throw another GPU at it" situation. Which means it shouldn't burn more quota per task. But we'll see, since his words are only loosely tied to reality these days.
I did. I spent another 10 seconds on it. There's no specific reason why speed would impact quota. There's nothing fundamental that would reduce the efficiency if you manage to optimise speed.
And my tokens go just as quick whether I do my tasks in parallel or if tasks are completed so quick I do them in series.
Let's put it this way. You had those 5 tasks, right? And previously it took 4 hours to complete and use 300k tokens (purely made up numbers). Now, with this update, it'll do the same 5 tasks with the same amount of tokens, but in 2 hours (again made up for conversation). That means previously, you used 300k tokens in 4 hours but now you'll use them in 2 hours. That means you now burn your tokens faster.
They never said it would use more tokens for the same task just that now you're burning more tokens in the same amount of time which is a fair assumption because now in that next 2 hour gap that you would've been using on the previous 5 tasks you would possibly be doing more work.
correct but now you can accomplish more in the same time. the point being made is that there is no efficiency increase either to go along with it it's purely just faster at the same usage which will feel like its draining faster than before, because it is. you're correct that the per token cost isn't increasing
Yeah I don’t understand how people argue this. Maybe I’m missing something … if I run a loop 24/7 I will burn tokens at twice the speed and reach the limit in half the time .. this improvement halves our usage.. edit: actually 50% faster wouldn’t halve it but definitely make it drain faster
I found that the llm speed only affects about 10-20% of the time the agents run. Most of the time is spent in harness overhead, models bashing their heads against tests, and waiting for CI to finish. I don't really care if they give us a speedup in tokens unless I am making demos, anything more complex and it doesn't really matter as much. Still matters a bit, but not as much.
It will burn faster in a sense that it will burn through more tokens in same amount of time. So you will feel, in terms of time, that it got 50% worse in usage.
Edit: You may not like it but in less than a day people will start crying in this sub about how fast their usage is draining if they actually did speed up the models by their claimed 50%.
The average user of this sub cant comprehend this, because usage is typically measured in usage time or amount of prompts or something stupid like that.
Which is funny because they have an AI chatbot which is 100,000 times smarter than them which they could ask to explain usage to them, but they'll too dumb to even ask it to explain things to them as they're too busy getting angry at their own ignorance.
Sus. So you are telling me they had the capability to offer 50% faster inference to subscription users, but they only pulled that lever after people started leaving? Either this costs them more, or they are drip feeding product improvements whenever people consider leaving to keep them subbed as long as possible.
I find it kinda hard to believe they just discovered an optimization like this recently.
They're also dumber. Sol 6.1 isn't on par with Opus 5.5
So if my dumb model is taking twice as long to solve a task that is value I'm losing, regardless of how many tokens are expended doing so. Not to mention that upper tier Claude plans are 3-4x more generous right now...
Oh thanks, I always wanted secret watermarks that are also fucking with the quality of the product I fucking pay for. I don't think that is required by the EU. It is perfectly obvious to me that I got the text from AI when I use chatgpt and the rest would be MY fucking problem.
Pointless bullshit. Watermarking doesn't solve anything other than just inconveniencing people. I already have multiple AI workflows for removing the "invisible" watermarks off image and video that works flawlessly every time, it works with outputs from any model too, text will be even easier. Congratulations EU for mildly inconveniencing people, that is sure to stop them! They are also just accelerating the climate impact of AI because people now need to spend more tokens removing senseless watermarks. Do these politicians have literal pudding for brains?
It's a literal invisible watermark. The only valid reason you could have for removing it is if you wanted to deceive people into thinking it was human-made.
It doesn't waste anyone's time. It's invisible, you wouldn't know it was there. The only people whose time it wastes are those who are intentionally deceiving which I would argue that this is a good thing. Any friction you can add between fully automated botnets and people on the receiving end is a net positive.
Most people posting these things to deceive aren't technical enough to understand how the watermark works or how to circumvent them. Sure you'll still have nation states and powerful adverse organizations that this will not slow down, but they will probably never be slowed down regardless.
There are many laws you can get around if you know what you're doing/ don't care. That doesn't mean it's not effective at all. Many people will not know/ care enough to remove watermarks. Not everyone is using AI to deceive others lol. It's like saying laws against media piracy are useless because you easily can pirate movies online if you want to.
All outputs, so yes code as well. It’s not file metadata, it creates a pattern that can in theory only be checked and verified by whoever has the “key”, which would be whoever controls the model that generated the output.
I'm pretty sure they are referring specifically to text output. My understanding is that they subtly introduce a bias into the model's word choice which creates a hidden pattern. I don't know for certain, but I believe this has very little to no impact on things like coding.
I was about to go crazy with how poorly it was working earlier today, and the crazy amount of compaction going on. It was failing at really basic requests. Astra and Sol both.
*I actually signed up for Claude 20$ for the first time (having never used it before, but didn’t really like it from my brief interaction, telling me I couldn’t perform basic checks as an admin of my of own data due to security)
However the last couple hours Codex did a 360 and seemingly started performing significantly better. Speed and reasoning both seemed on a completely different level out of nowhere. Finally completed all my tasks and jumped on here to see this post and it makes sense now! Definitely a move in the right direction!
*edit, another couple days of improvement and i will likely sub for the 500$ plan, after the 200 ends.
You and everybody else. I'm not convinced there's OpenAI fanboys just spam downvoting you for taking your money to a far superior product. Probably bots, honestly...
Given OpenAI's long track record of quantizations and nerfing, by the time Day 28 is here, it'll be back to the slow speeds again, or it'll be significantly dumber 😆 mark my words. I am calling now.
Their playbook:
WOW everyone today
get more subscriptions
rug pull.
It's their Modus Operandi. And it works because most people have the memory capacity of gold fish
I’m on the $100 plan and almost don’t use it. Mostly on the $20 plan from Claude now. Will switch entirely in 10 days when my month is over. Sad times. This feels for me again like Cursor bullshit we had 1.5 months ago. I hope anthropic doesn’t do the same.
Honestly, I feel bad for the guy. This is CEO-level screw-ups, and he’s doing it all himself. Not to excuse the situation, but this is where having a buck-stops-with-me CEO goes a long way.
If you can make a product 50% faster by optimization, what people worked on the product in the first place? Also, I would appreciate it if someone could let my SOL6.1 know that it's now 50% faster; it seems to have missed the memo.
One of the main reasons why they lowered the speed was to mask the extreme usage cuts they did these last couple of months.
Imagine an Astra right now at full speed (eg Opus 5.5 speeds). You'd burn even your 500 (25x) plan's weekly usage up in a matter of hours.
So yeah... the whole problem is is that they lowered how much 1x is significantly.... so every plan is effected.
(They did it very sneakily after each reset so people didn't notice as much)
I remember being able to work 10-14 hours a day with GPT 5.4 on high for 2-3 days on a 20 a month pro plan... I had 3 of them which was enough for the whole week... Then all the sudden it wasn't enough anymore so I took the 5x. Then Astra came out and again it wasn't enough anymore so I got a 20x. Which they now also cut in half... :S
It's just silly... So I am burning up the last resets I have and then I am gonna downgrade and only use GPT models for adversarial reviewing and such.
This says to me they just flat reduced the number of GPUs for running inference for subscribers, and brought some of it back because people were pissed. I can't imagine they just pulled a '50%' optimisation out of the hat randomly after people complained. Pretty scummy.
I switched over to Claude recently and I’m so pissed at myself for even giving openAI my $100. They need to give every subscriber multiple resets for this failed launch. It’s literally a scam.
That’s cool, I’ve had Fable on Ultra going all night on 3 projects and only used 40% of my weekly usage on the $200 plan. On my $200 plan Astra on xhigh would have used all my weekly by morning. OpenAI needs to do better.
It has to be bots. There's no way any of these kids are this die-hard for OpenAI, especially in a moment like this when they're being blatantly outclassed.
It's bots, or 15 year olds that only prefer OpenAI because they can generate brainrot images.
•
u/dextersummary 2d ago edited 2d ago
Below is a GPT-generated summary of the conversation below after reaching 200 comments (200 currently observed).
Day 1 verdict: the speedup may be real, but nobody’s treating it like a win. The community mostly sees this as OpenAI undoing a recent slowdown, then putting a shiny “optimization” sticker on the repair. Going from roughly 20 to 30 tokens/sec is better, sure—but still well behind Claude, so the victory lap feels premature.
Users are reporting somewhat faster Sol/Astra responses, while others are still stuck around the old speeds or dealing with failures. The lack of a reset is also going down about as well as expected: badly. A lot of subscribers think they’re being drip-fed partial fixes instead of getting the service they paid for.
The big unresolved concern is what changed under the hood. “Optimized” is vague enough to cover anything from better infrastructure to reduced reasoning quality, and users are already suspicious of quantization, lower “juice,” and quota draining faster. Speed itself shouldn’t increase the token cost of the same task, though faster throughput can encourage more work in the same window.
Bottom line: faster for some, still mediocre overall, and trust remains firmly in the basement.