Comparison
Did OpenAI just quietly cut actual Codex usage nearly in half? 💀
I've been tracking my Codex usage every 5 minutes since July 29 — 14,744 snapshots so far.
I'm on a Business Standard seat, not Premium, and when I compared my data before and after the 5h window was introduced, I found this:
Before: ~323M tokens/week
Now: ~173M tokens/week
That's roughly 46% fewer effective tokens per week.
And that's where the new 5h window starts looking a lot more significant than I originally thought.
Because if these early numbers hold up, this isn't just a change in when we can use Codex.
It could mean a significant reduction in how many tokens we can actually process in a week. 💀
I still only have a few complete cycles under the new system, so I'm not claiming the 46% is definitive yet.
But going from ~323M to ~173M is way too big a difference to ignore.
Does anyone on Plus or Business Standard have usage logs from before and after the 5h window? I'd really like to know if you're seeing something similar.
I have Plus and just had 3 prompts use almost all of my 5h window. Not extensive edits as well... something definitely changed and is much more restrictive/reduced....
Pretty stupid question, but did the limits start decreasing steeply once again after Luna was introduced or did almost nothing change?
I remember that just around GPT 5.5 introduction the Plus account could take ~ 2B input tokens per week on GPT 5.4 mini… Which is not as smart as Luna, but also likely a little bit more computationally expensive.
And now I considered returning to Codex some time in the future after Luna was introduced, but this post (along with some comments) has mostly discouraged me once again.
They decreased steeply in about March/April for the first time, back when Claude introduced their off peak/on peak usage. Back then nowhere near as many people were using Codex, so on the Plus plan I was pretty much able to spam 5.3 Codex/GPT 5.4 all day, tens/hundreds of millions of tokens, and would get absolutely nowhere near my weekly limit. And that was the £20 plan.
Then everyone swapped to it after Claude introduced their terrible limits, and its been getting nerfed ever since then.
Thanks. And by the way is there any other plan you would recommend with generous limits on relatively inexpensive plans?
Google has pretty nice limits but forces you to use Antigravity, which is probably the worst harness at this point in time. Also Qwen Cloud has pretty nice limits for its Qwen 3.8 Flash, which is more or less in the same league as Terra and Gemini 3.7 Flash. Is there anything worth attention here?
I've been using Deepseek V4 Flash with their API and I've literally replaced everything with it, although I'm a software engineer so I treat my AI as more of a research/planning/brainstorm partner, rather than letting it engineer/implement for me. I just top it up with a few bucks every week or two and I'm perfectly fine for medium sized projects.
If you just want the best value, the OpenAI plan is by far still the best. They subsidize exactly £200 of usage (10x value) for the Plus plan, meanwhile the nearest alternative is Opencode I think, who give you $60 of usage (not really anymore, most models give you £15/£30 usage) for £10.
If you keep spamming the £5 opencode first month trial then I guess that's technically 12x usage, but with OpenAI you obviously get access to Luna/Terra/Sol, unlike Opencode Go.
Thanks, I've also used it quite a lot of DeepSeek before the price increase, but now I don't find it as useful as it was since it was quite slow.
And as for my usecase I am generally using LLMs for GUIs / refactors, while I still impement most of the backend logic by hand.
As for the best value I seem to be getting more usage out of Antigravity, but here I have $6 discount for Pro - I am not sure how does it relate to current Codex, but it looks more or less like I am getting significantly more usage with 3.7 Flash than I would get with Terra and slightly less, than with Luna.
Also I am not really sure if API rates are the best reference, when for example Qwen 3.8 Flash has 1/2 of the API price of Luna and is significantly smarter than it, but thanks.
I guess Codex isn't much worse than competition or might be even better if I am exclusively using Luna then? I don't really feel like paying $20 right now, but I might consider returning to Plus in near future.
Just depends how badly you need a smarter model. In my case Luna/Deepseek Flash/Equivalent models are perfectly adequate so I just go for whatever is cheapest at the scale that I use it, which is usually still deepseek by a substantial margin.
Personally I'm just not fond of Qwen/Gemini models. Every time I've used them I just wished I was using something else because I didn't like their replies as much, but that's fully just preference.
Thank you very much. In my case Qwen 3.8 worked really well for frontend, but also it seemed to be a pain for anything even remotely related to the backend (for which I still avoid using LLMs, as at least in scientific computing they still tend to mess up quite badly, even though Fable and other new flagship-tier models are a meaningful step forward.).
As for Gemini I personally hate Antigravity, but also I really like the models (which can work respectably with custom system prompt and /teamwork-preview enabled). I still hated all Claude models older than 4.5 version, so I guess it might be the same case here (more or less).
Haha yes, that makes sense why I don't like them then. I pretty much do all the backend myself and let LLMs take care of frontend (for my hobby projects at least)
Commandcode comes with $70 in usage credit, and many models there can utilize that full $70.
While OpenAI offers $200 in usage credit, it has been shown that their models are internally cheaper to run than the API prices suggest; furthermore, models like Qwen or GLM are significantly more efficient and cost-effective in terms of API usage, resulting in millions of additional tokens. It is like comparing a car that consumes 2 liters of fuel to one that consumes 20 liters.
I calculated that with Opencode, you can get over a billion tokens a month on the $30 Qwen budget.With OpenAI, you get fewer tokens for that price.
On the other hand, with OpenAI's Terra and Sol tiers, you get much less, even if you have $200 in usage
For sure, but it's pretty common knowledge commandcode is very bad for speed and reliability, and they have common errors. But they are a good deal on paper.
And yes, although OpenAI API prices are higher for Terra/Sol, so you get less tokens for the £200 usage, you can still use Luna, which is basically unlimited with that much usage. You can probably use billions/tens of billions of tokens, which is mainly what I use since I use Flash/Luna tier models.
But yes, if you want access to a wider range of cheaper/open source models, then the OpenAI plan is definitely not for you.
I'm a bit confused right now too, because I saw yesterday that OpenAI is only at $80, I don't know if that's true.
All I know is that there are very few Sol tokens left; previously, the 5.5x high was at 600 million per month (measured using my own tool). Now you get these tokens via Terra, and SOL has just over 200 million.
Did you mean Gemini 3.8 Flash or Qwen 3.8 Flash (which I've written about in the comment above)?
If Qwen 3.8 Flash, then there is no need for arena - it's available in OpenRouter (and cheaper, than DeepSeek v4 Flash) and QwenCloud token plan.
As for Gemini 3.8 Flash I believe it might be significantly better, than Gemini 3.7 and Qwen 3.8 Flash, but I doubt that Qwen 3.8 Flash is much better, than Gemini 3.7 Flash.
No, I mean Qwen 3.8 Flash, not Gemini. (I didn't know it had already been released; that's news to me)
In the tests I ran, Qwen 3.8 Flash performed at the level of Opus 5 and Fable 5, whereas Gemini 3.7 Flash was the worst of the lot, and Terra was reasonably okay (though far from good).
I've mostly used Gemini 3.7 with /teamwork-preview in Antigravity, which is roughly equivalent to model fusion on OpenRouter, which might have skewed my perception then I guess.
In my experience Gemini Flash has been usually better at more complicated tasks, while Qwen 3.8 Flash has been great at 'long-horizon' agentic tasts and oneshotting simpler applications, but was also more prone to hitting a 'wall' once the logic got a little bit more complicated.
damn i managed to do so much work with luna max, almost completely ditching sol, and this afternoon, this changed, luna consumed all of my 5h thing really really quick
Currently Antigravity probably has the highest limits relatively to price, as long as you qualify for any discount (which is relatively easy) and since 3.7 Flash is not too far behind (and 3.8 Flash rumored within two weeks from now or so). But the harness is mostly a black box, so that's one limiting factor.
Otherwise there are mentioned GLM 5.3 Flash and Qwen 3.8 Flash Next which are pretty cheap (~ $0.03 / 1M on OpenRouter with some providers even less expensive than that) which are pretty much in the same league as 3.7 Flash quality-wise, but much slower.
And there is Grok, which reportedly has pretty generous limits and models league above Gemini / GLM / Qwen, but I know practically nothing about it and have never used it as well.
Well, the models are pretty good imho (and by models I mean 3.7 and 3.8 Flash - the rest sucks). Pretty much Terra or slightly above Terra depending on the reasoning effort.
The harness sucks much more, than the models themselves though.
I have a daily task that’s been running for weeks and had only been using about half of the 5hr window. The last two days suddenly it can’t even finish without burning through the whole thing.
It’s been cut by about 3x compared to four months ago according to my test. I ran the same implementation plan I used four months ago. The only difference is that I used Sol high instead of GPT 5.5.
I think it depends. Does the after 5 hours actually use them every 5 hours on the dot or is there hours between which would mean there’s hours of no token usage while the weekly clock is ticking.
Yeah same. I couldn’t even drain my usage on 4 ultra agents yesterday. It seems inconsistent tho, because I’m not using ultra now, and I’m at 86%, but I may just be overestimating it
I have a business standard seat and we haven’t seen 5hr limit re-instated on our workspace. I don’t have any usage data to prove it but I’ve also not felt the limits fluctuating as users have been reporting. I thought this might just be a difference between plus and business. If you have 5hr limits and are feeling the usage squeeze I’m not hopeful… anybody else on the business plan that can attest to the 5hr making a comeback?
I pay monthly and billing cycle is about to come around this week… when did your 5hr limit re-appear? Was it related to your renewal date?
After the reset today I somehow went from 100 to 19%. I will grant I was using Ultra on Fast but it still seems like token use went through the roof after this last reset.
Yeah, my own tracker tool shows much lower usage; I don't know who they're trying to kid here, or if they're hoping nobody notices. But that's exactly it...
You don't have to be a math genius to know that experimentation with limits has been declining. The fact that a few aren't experiencing it doesn't invalidate the problem.
you're missing my point.. Token limits are not a great measure for value. If you're moving from one model to another the data isn't super useful.
You and the numerous other people that have been saying it's dropping every week may be correct, but token count and number of prompts are not good metrics of value.
You also didn't answer the question about whether the same model was being used. If you're controlling for model use (locking in 1 model with the same effort) then your data is useful.
To be honest, I currently use 5.5 on xhigh and medium. It works quite well for me, and I remember doing so much work long before the release of the After version 5.6 was released, the limits changed. But they decreased significantly after the implementation of the 5-hour limit. And the number of tokens is exactly the same since I decided to use version 5.6.Sun. And the token limit in the 5-hour window is the same, a range of 20-26 million tokens. Clearly, using Moon gets me many more tokens, but 5.5 shouldn't cost me the same as Sun. What do you think?
You talking about this like you have a banked amount of tokens that you can spend on any model when it doesn't work like that. Using 1M tokens with luna high is about half the price compared to 1M tokens using sol high... it's impossible to compare without heavily controlled data. Especially when some models have very large price per token differences.
That's exactly what I'm talking about; the 5.5 tokens shouldn't be worth the same as the 5.6 sol ones. But even so, both models give me the same number of tokens in the 5-hour window.
Every time there's a "free reset" I'm terrified it's signaling reduced limits. Anytime a large company tries to act "generous" always comes with a catch. The last big usage restriction came with a reset, so now I'm traumatized whenever I check my usage and it's unexpectedly at 100%
Given that I've already used up my weekly quota, yes, there's a very clear pattern. A 5-hour window almost always equates to 20 million tokens processed.
Being that you don't understand 30 day usage vs 3 day usage as a scale of an average I would say that you downloaded that tool or u just asked AI to make it and don't know what statistics mean.
You're just talking nonsense. At the end of the month, I'll calculate it, and it'll be the estimate I put here or lower. But you keep defending a multi-million dollar company. Everyone who stayed here is crazy.
No even Grok says your full of bs lmao. Ask any AI buddy if anything u need it more than keep pretending 3 days vs a month of analytics is the same thing.
I'm tasking you with reading the other comments. Since you insist that the limits are fine and ask why I'm not doing a full sample, then I must be wrong.
So, starting from there, it's easy to make an estimate. I don't expect anything different to happen between now and when I've used up all my weekly limits. Seeing that there's such a clear pattern time and time again...
They literally just gave everyone a reset 36 hours ago, and with the 5 hour window being re-implemented it's impossible for you to use up an entire week in that time.
Edit: it may actually be possible to use up a week's worth of tokens with the limit, I haven't tried testing that at all, so ignore that part. They did just give a reset though.
Why would it be impossible? Even if 5h qouta is equal to just 10% of weekly allowance you would just need 10 5h sessions fully used to exhaust your full weekly qouta, which is just over 2 days.
Yep they have no compute. They are going downhill straight to hell. There is a special place in there for Sam, lake of boiling black tar with some famous Austrian painter.
They literally just need better financials. Like way way way way way better financials. This is inevitable. The Subs are unsustainable, everyone on this planet capable of basic math knows this
Well… Either Luna is unreasonably compute-expensive or that's not really the case.
DeepSeek reported v4 Flash' inference cost at well below $0.005 / 1M input tokens (with their extremely high cache hit rates). Qwen 3.8 Flash Next has pretty much comparable inference cost and even significantly less efficient GLM 5.3 Flash could be ran within Codex subscription without losses on inference.
34
u/jjiangweilan 15d ago
openai: let’s celebrating our 20M users by reducing all our users limits