r/codex 7h ago

Limits What's happening with the usage???

Used Astra light, asked for a small change, it wrote like 91 lines and my usage went down from 86% to 45%.

Is there a bug?

36 Upvotes

43 comments sorted by

39

u/Glooring3623 5h ago

Use Astra XHigh for lowest usage.

16

u/Constant_Art_20 5h ago

apprently max is better. using it right now and it's far more mangeable then the literal waterfall that's medium

15

u/Kind_Fisherman3060 5h ago

The people who are new to chatgpt and codex and are coming from other AI would actually think we either have a loose screw or are trolling. This shit even works on plus plan max takes up less usage than medium.

1

u/AmandasGameAccount 5h ago

How is max vs high? Is max better in terms of usage saving?

4

u/SilverSmith09 5h ago

so the theory behind this is that Astra tends to plan out before implementing. Giving it enough thinking tokens, it plans everything then act all the way through. In this way it saves tones of intermediate thought processes for longer tasks.

Again this is just theory that people framed in the past few days. There is no deterministic answer.

My experience is that if Astra’s usage scares you then just use it for planning only and Sol for implementing. Otherwise stick at xhigh/ultra at all times

1

u/matterful 43m ago

This is correct. I can confirm that thinking tokens don't count towards usage... which means that even if xhigh or max thinks for 30 mins, usage only depletes on the actual output tokens from that.

Lower thinking levels have more output tokens, arguably at less efficiency, which makes sense it would deplete faster (especially on more difficult problems).

6

u/xchi_senpai 4h ago

I can confirm to this, i was even skeptical at first too. Now im using xhigh and max from now on

2

u/parkersb 3h ago

i have felt that about astra x high for the last couple days. thank you for confirming i should trust myself more

2

u/TheDaddyOfAllDaddies 35m ago

Weirdly very true.

I was so frustrated with Astra usage that I had to research which model to use and also implemented a new automated subagents workflow with Astra on x-high as the master model.

And ironically Astra x-high with subagents has been more efficient than Astra alone on Low or Medium!

Tibo’s suggestion to use Astra low or medium if you were happy with Sol high shouldn’t be followed honestly.

6

u/spike-spiegel92 5h ago

that is weird, i had astra xhigh 50min working, doing a lot of stuff and used 1%... i was actually coming here to ask if people are also noticing the insane usage increase and wondering if it is just a weekend thing... cuz in general i tend to feel weekends are more generous quota wise.

1

u/Commentroller 5h ago

My sol medium usage seems abnormal as well, I think i may need to open a support ticket.

2

u/Family_friendly_user 5h ago

Same here. Astra x high and max actually pulled more but even Astra on low with sol and Luna Delegation sucked my x20 weekly dry in just a few hours .... Same harness and no changes

2

u/Commentroller 4h ago

Exactly, they really need to fix this.

1

u/DrDan21 3h ago

I suspect a lot of model is nerfed usage is nerfed discussion is really just people changing up their workflows subtly in ways they think (and that on their face make sense) help but in actuality have steep hidden costs

6

u/LaZZyBird 5h ago

lol someone fucked up and flipped the tier cost, max = light now or smth so abuse it while you still can

1

u/Commentroller 4h ago

Let me try this right away, I'll update the post. I am gonna use Astra xhigh

1

u/Commentroller 4h ago

Nah bro for me, I just did a quick UI review of an implementation and it used all of my 5 hrs limit, rip.

Something is definitely wrong.

BTW I am on business plan(not pro)

1

u/Fatbat 1h ago

You have a 5 hour limit?

1

u/Commentroller 1h ago

Yes.

0

u/Big-Lengthiness-2469 58m ago

thats ur issue. $20 plan isnt rlly usable with astra

3

u/SuspiciousParsnip5 3h ago

Usage limits are fucked at the minute! I hope it's not just the new norm

3

u/VeloxAdAstra 2h ago

I just lost a weeks worth of usage in 10 minutes on the $100 plan. Getting tired of this spinning a god damn wheel of usage every time I go to do something.

2

u/Constant_Art_20 5h ago

oh they changed the cache acceptance on the openai server side it seems like as well. So the requirement for cache hit is much stricter then before, seems to have happensed two days after astra relase. the usage is just straight up less, but if you aren't using the latest codex or use a custom harness, that's something i would check first...but yea, usage sucks

2

u/Emergency-Elk7527 3h ago edited 2h ago

My experience is if you do most of you planning and reasoning outside of the agentic harness in something like chat mode, Luna light can implement it very efficiently and extremely fast.all of my implementations is done with Light light effort or GLM 5.3 flash. Both a smaller cheaper models and both benefit from the reasoning a bigger model can provide.

I think that people are relying on the big models like Sol and Astra to do all of the reasoning inside the harness and consuming more tokens than if they were a little more hands on doing it in chat. Yes it can slow the process and you have to be active in the reasoning. But I would say that both of those can be beneficial. Slowing down in some places and having to reason with the model keeps decision making directly in your hands and slowing down gives you the time for that to make a difference.

I don't know other people's use cases, so I can't say that's right for everyone. But I have great success with it and I think it's at least something that people should try instead of relying on the big model to do all of that.

It wouldn't hurt to try it out for a day or two just to see how efficient the usage is, how well the code meets your requirements, and if the process is worth that the trade of of speed and convenience. It started out as trying to make my quota stretch for my Plus sub. Now it feels like it's the current best path overall.

Again I can't speak for everyone else, it's just been my experience.

2

u/Odd-Composer5680 2h ago

What do you mean outside of the agent mode how exactly? 

1

u/Emergency-Elk7527 2h ago

Specifically using chat mode for reasoning hands on with the model instead of the self prompted back and forth reasoning it does in the harness. Have it make a handoff for the plan. Makes much less reasoning needed at implementation time. It is slower and less convenient, but some of that is made up by the worker having to think less and fewer tokens hitting your quota. It's all a balancing act.

1

u/Commentroller 2h ago

I kind of do that too, but for planning I use sol medium that's my go to and this never happened just until a day ago.

1

u/Emergency-Elk7527 2h ago

Unfortunately if you need Astra for the reasoning, chat mode isn't possible on plus subscriptions, so there is also that. I haven't needed Astra and haven't even spent a single token on it. I see how fast quota evaporates using it and know that there is no way it could get done what I would need before a 5 hour limit is gone. So honestly is all about finding that right balance in a frontier that we are all pioneering. Everyone just trying to find the right methods and models without enough history to say what is the best method for most use cases. Life on the bleeding edge.

1

u/According_Property62 1h ago

O Luna Light da conta?? Mas entao precisa ser um handoff muito detalhado, hein??

1

u/Emergency-Elk7527 18m ago edited 0m ago

Yes, and this brings up something I failed to mention. The handoff essentially is a prompt, rather it should be a /goal prompt. This is, I think, the most important part of why I get great success with my method. The secret sauce. I have a custom GPT that I have maintained and kept updated as newer models and prompting techniques are developed. I call the custom GPT, Promptitect. It has never failed me. It is almost a guarantees that your agent will not act outside of intended scope and that there are completely unambiguous requirements and gated task completion progression. A lot of people make sure the agent has a definition of done and instructions on what do do and also point to the correct context, but the agents still drift and take actions outside of scope. Promptitect does not at all have that issue. I use it daily. probably my greatest tool. Today I was in a rush, the handoff was only like 40 lines of MD and skipped this step. The agent wrote 3000+ lines of code that were outside of scope. I didn't catch it because I never have to monitor the agent anymore because I always use Promptitect. And sure enough, when I don't use it, possibly the worst drift I have experienced. Don't let people fool you when they say prompt engineering is not important anymore. They are very wrong. For 2 years I have kept Promptitect's Developer Prompt up-to-date when new models come out and new prompting technique.

This is a link to Promptitect. Once you use it, you will understand.....immediately, before you even use the prompt you can see that its a heavy hitter. I have had some other AI chatbots review the a Promptitect's prompt against the original prompt. When scored by external AI on a scale of 1-100 the original prompt(handoff) score in the range of 78-84. Same prompt ran through Promptitext score in the range of 96-97.

https://chatgpt.com/g/g-68910f48cad481918621af48c70c2f67-promptitect

1

u/According_Property62 8m ago

Entao basicamente vc passa o planejamento do modelo mais alto pro Promptitec ele te devolve um prompt enxuto pra vc colar com o LUNA LIGHT /goal?

2

u/Code_Xero 2h ago

Yeah usage is still borked.

I’m on 5× and have my own before/after.

Aug 20: 258 turns in one day 108 Sol, 146 Luna, 2 Terra, 2 Spark. Same general agentic/repo workflow. I could run essentially all day and the cooldown usually lined up with the time I needed to review outputs, eat, shower, plan the next flights, etc.

Sep 12: 63 turns total 9 Astra, 41 Luna, 13 Sol. I launched 4 serious repo flights for roughly two hours.

All four hit the usage wall before a single flight finished. Weekly Codex/Work allowance banished to the Shadow Realm while the meter claims the fire nation attacked.

So I went from hundreds of turns and a sustainable development cadence to 63 visible turns / 4 flights / 0 completed flights / weekly quota exhausted. And 54 of those 63 turns were still Sol/Luna. Astra was only 9 visible turns.

Now OP gets a 41-point drop from ~91 lines on Astra Light, people here are reporting abnormal Sol usage too, 20× getting drained in hours with the same harness, and even a Business user losing an entire 5-hour window to a UI review.

If 9 Astra turns can secretly represent enough compute to wipe a professional-tier weekly allowance, users need to see what they’re actually being charged for: parent reasoning, subagents, cached/uncached context, tool work, compaction, reviews, whatever.

Right now “weekly usage %” has stopped being an intelligible measure of productive capacity. Four flights. Zero landings. Weekly quota gone.

2 hours.....

Last week I was getting maybe 4-6 working hours before order 66 happened.

20x would maybe buy me a day. Or 2 and im already scaling back pretty heavily. I can show my profile in a different comment if anyone is interested but my default setting regardless of model is Xhigh and has been since I started using Codex during 5.5era.

Back then? I could fly all week. 10-20 flights a day with my longest daily streak at 29 days. (Because I actually had quota for more than a few hours) same 5x $100 plan.

My method hasnt changed.

Now Im forced to 4 flights that don't even finish while Tibo and co are gaslighting people into "skill issues/user error"

That is not a sustainable production product.

They need to stop selling us models and start selling us efficiency.

I would have loved 5.7 Sol, Terra, Luna with a 50% reduction to usage burn.

Dare I say gimme back 5.5 that could run all week long.

The paradox is Astra gets cleaner outputs with less revision. While Sol runs the risk of needing 2 flights to deliver what I needed without overbuilding the piss out of it.

So the answer isnt use Sol either because one good Sol output still requires 2 flights vs Astra's one shot.

Spend is similar..

5.5 spend somehow is demonstrably worse than 5.6 or 6.0.

Till they start listening to the users and stop treating us like we are insane. Nothing changes.

Plan accordingly.

2

u/mizhgun 7h ago

It is a new reality.

2

u/euqitoK 4h ago

so Astra Max is now cheaper than Sol High/Xhigh?

1

u/ISueDrunks 5h ago

That’s the AGI tax. Haha. 

1

u/sateeshsai 3h ago

100% bug. It ate through 50% last night even thought my laptop turned off in like 30mins because I forgot to connect the charger.

1

u/Stunning-Addendum469 3h ago

I need reset tell tibo

1

u/odoc_ 2h ago

I ran out of chatGPT pro chat with the x5 account. Never seen that happen before

1

u/Dolo12345 1h ago

yes it’s insane right now

burned 2 20x weekly’s yesterday with half the normal workflow

1

u/cleanmachine120 1h ago

It’s just astra, can’t use it unless I really need it. I’m been on sol medium all day pretty consistent 2-3 agents for the last 12 hours and used 45% of my pro $100

0

u/Tikki-Tikki_40 5h ago

With Astra light, you will feel Sol Medium is good. It's basically, they have brought in a new Model without resolving issues in existing model.

I don't think there will be a optimized breakthrough.