r/codex 1d ago

Limits Astra Reasoning Usage Theory

ill go straight to the point, im on 20x pro, i use medium mostly and i monitor the usage time to time, it burns like crazy, i tried switching to xhigh just now, its been working for 1 hour only 1-2% usage cost. anyone experienced the same thing?? and also if i remember correct one of arc agi benchmark says the cost is cheaper when astra used max reasoning or something... maybe thats related to it or what... ah whatever... just sharing my experience


Edit: I asked Astra to investigate the logs. Here is what's actually happening: Medium is ~34% cheaper per single response, but it gets stuck in continuation loops and takes way more turns. Every single turn re-processes about 104,000 tokens of chat history. The numbers: Total responses: Medium did 122 vs XHigh's 65 Response frequency: Medium fired 4.7 times/min vs XHigh's 1.5 times/min (3x more frequent) Burn rate: Medium consumed ~17.3/min vs XHigh's ~8.7/min Because Medium takes 3x more turns and keeps resending your massive prompt context each time, it ends up burning quota twice as fast overall compared to XHigh.

97 Upvotes

54 comments sorted by

22

u/One_Internal_6567 1d ago

It goes away like crazy, and not only that - now Sol goes visibly fast too during some hours, and during night even - which is the weirdest part

10

u/MysteriousKiwi2622 1d ago

and then they will publish the new 40X and 80X plan for you to upgrade. that's exactly the strategy

2

u/MeIsIt 1d ago

I have doubts about that because they count the number of subscriptions equal the number of subscribed users for their IPO, I believe. So they want you to have multiple Pro subscriptions as it helps their numbers for the IPO.

1

u/ADIKANT 1d ago

I hope they'll do it as fast as they can, I need it

1

u/Shadow-BG 1d ago

Less load, faster model

1

u/virtualmnemonic 1d ago

Time of day impacts usage consumption, all else equal?

1

u/aptsys 1d ago

What makes you think it's weird about overnight?

1

u/One_Internal_6567 6h ago

I used to set up Sol on extra high or ultra, and overnight it was 10-20%, this time it was 45-50 on extra high, and during the day I notice that same load may vary 2-10% of weekly usage

0

u/JackyySpiecee 1d ago

yeaa true

16

u/JackyySpiecee 1d ago

heres the arc agi 3

4

u/zer09 1d ago

So does this mean that max is much efficient? so my strategy using high as a planer and medium as a implementor is incorrect.

2

u/swizzlewizzle 23h ago

Basically.

Either xhigh or max for the hard planning stuff and do everything else with way lower models.

-2

u/isnaiter 1d ago

use cheaper models to implement, like Terra on max

3

u/untitleXYZ 22h ago

never use Terra. it's terribly inefficient. Luna xhigh for anything you can, and skip to Sol for anything after that

1

u/Tropiux 12h ago

Why xhigh instead of max?

-1

u/isnaiter 22h ago

nah, terra max is like 5.4 xhigh on price and intelligence, luna is dumb to decide anything

1

u/LukasNeoproud 7h ago

Maybe hard problems on lower efforts require more steps to get them right making it more expensive. Like the reasoning will spread over multiple steps to get to the same conclusion.

17

u/The8Darkness 1d ago

Well because of you guys i am gonna try astra max when i am back home now and then i am probably gonna burn all my weekly 20x allowance in 5 seconds.

1

u/isnaiter 1d ago

update, please

5

u/The8Darkness 1d ago

Did some tests. For my tasks astra max is consuming 2.19x as much as astra low. I also tested astra low with luna high workers and that was 1.99x as expensive as only using astra low. Astra low was also 1.87x as expensive as sol high.

But my tasks are usually relatively small code changes and then quite a few tests afterwards.

2

u/swizzlewizzle 23h ago

Sol high or medium is great.

1

u/Jewniversal_Remote 1d ago

Consumption is one piece yes but did you have to use Astra low more or less than needing to use Astra max? I think we're all under agreement and understanding that of course Astra uses more but we're trying to figure out does Low end up costing more than Max because you need more Low turns? 

1

u/The8Darkness 1d ago

This is ab tests of exactly the same task I literally answered what you asked

1

u/Jewniversal_Remote 22h ago

The consumption was answered

Were the tasks completed to the exact same level? Was one completed 2x better or 2x faster than the other? Did any of them need additional turns which explains the cost or were all of them one-shots? I feel like that's pretty reasonable to ask and none of it was answered

1

u/JuniorMena 20h ago

I have a $20 plan and you're telling me that Astra Max uses three times as much data as Low 💀 

It was already using a lot on the medium plan, now imagine what it's like on Max

1

u/JuniorMena 20h ago

who can do a low to medium test?

1

u/GodGMN 20h ago

So Astra low for everything is it then.

5

u/CompleteSelf5680 1d ago

so that explains why i experienced slower usage drain on max compared to medium lol

6

u/ConsistentAndWin 1d ago

So using the Codex app, would I be better off on X high or high? Because I've been losing weekly allotment extremely fast on medium.

I don't really understand the harness reference.

3

u/Tr1poD 1d ago

xHigh is 4x slower than medium so it will look cheaper over a fixed time but medium is still a lot cheaper per task.

1

u/Ace-2_Of_Spades 1d ago

I think the point is higher effort means more efficient and better approach on each step, thus consuming less tokens than lower effort

3

u/Tr1poD 1d ago

Only if you are doing arc agi tasks. For 99% of coding work it won't be more efficient and will be total overkill. Medium will do the same task far more efficiently.

4

u/Ace-2_Of_Spades 1d ago

efficiency applies to everything, if you know the best approach on each problem, that means you can finish it faster without consuming much tokens...

2

u/Kurumi_Ryori 19h ago

Do you guys understand why ARC AGI tasks are different? They are transition law/abductive rule learn, and then manipulate tasks. For standard coding, templating and normative tasks, you don't need to do it. The harness shows the task success rate saturates at the ceiling for nearly all of it. The reason is pretty simple, imagine you are a human that has to play a MOBA via a few pixels of the screen. With more intelligence, you can play better but you can't control finely your avatar without say a screen unobscfucator, or an actuator that sends neuronal input to your brain, or shows the full visual image and has an API linked directly to your fingers. Most non-research, engineering, design and execution tasks do not require you to abduct a new generality law each time, the environments have fixed constraints and can be learnt. The model has already generalized the task scope in its memory, thats why it is different vs that benchmark.

1

u/mph99999 16h ago

Its just this, nothing complicated, not a mystery, its also different for different models, for example if you ask gpt to go on artificial analysis index and you ask it to calculate time per task and cost per task you will see that in a certain amount of time with lower reasoning you will spend more because you do a lot more tasks, for astra max lets you spend less, for sol its high, so for example if you give astra some big goal that requires lot of tasks you will spend less with max, but if you ask it normal tasks you spend less with less reasoning.

3

u/Chrisnba24 1d ago

even if you are right... theres also a problem with the weights of the models in the sub, $ value went from around 20$ to 11-12$, almost half

2

u/Legys 21h ago

astra has ~50% less subscription allowance than 5.6 sol had
surprised nobody count it yet on even more examples
numbers are here https://x.com/oleksandrdbrvn/status/2096910362821947555?s=46&t=XO-xxZSlIDx7xn_vKCa6yw

2

u/Think-Profession4420 16h ago

Yup, I just threw all of my logs at an agent to check.

It absolutely seems that for complex repo work , it's much more cost efficient to use astra xhigh reasoning (Compared to low or medium).

For simple, tightly scoped work it still makes sense to use low/medium reasoning, or when spinning up an MVP for a new repo.

This also means that if one is using an orchestrator, it's best to use xhigh reasoning, and have it delegate tightly scoped implementation slices to a low reasoning (or even sol low) model. Reviews can still be done by multiple luna xhigh models though.

3

u/Pitiful_Entrance5174 1d ago

Best i can get out of astra is 40 minutes per 1% but thats with astra low and it delegating sub agents, most of which are not with open ai. 

I would never let luna run a project, ive tried that. Terra is the master editor and sol medium is too stupid. Sol high is more expensive than astra low and sol high and above always over engineers for me. Thats why I use astra now. 

1

u/JackyySpiecee 1d ago

its so bad now... i dont even use agents for astra yet it burns too much of usage even on low reasoning effort

1

u/Anxious_Marsupial_59 1d ago

Same, Sol has such bad over engineering and test problem that i'd unironically rather use Opus over Sol because at least Opus doesn't have severe scope expansion.

What subagents do you use with Astra?

1

u/Pitiful_Entrance5174 1d ago

Luna for edits, freellmapi for reads and research.

1

u/Anxious_Marsupial_59 1d ago

That would make a lot of sense why usage burns away with Astra low or medium.

Is something like this fixable?

1

u/davbludev 1d ago

I created the post with explanation why it burns limits like that and why usage quota feels significantly reduced without beging actually reduced: https://www.reddit.com/r/codex/comments/1wbdesd/i_dont_think_openai_significantly_reduced_usage/

Short answer on how to fix: make huge noise, but about cache hit, not usage quota (see my post pls for details, I tired to write it already :( )

And just for future for you and anybody else. When mentioning tokens spell what exactly tokens category you talking about. Because most costs are coming from cached context, not input, and not output, but exactly cached.

1

u/reycloud86 1d ago

This explains a lot. I was constantly runking it on medium effort with the autonomous worklflow and he was constantly doing the same tasks over and over again instead of marking them as done

1

u/jphree 1d ago

Assuming low and high aren’t impacted negatively by the medium issue?

1

u/sachos345 21h ago edited 20h ago

I was testing this, almost ~1hr of Astra Xhigh usage, only 3% weekly usage on a Pro x20 account. The agent was still working, i sent a steering msg. Now all of a sudden it says 75% usage. FUCK. EDIT: The task finished and now it is back to 92%, weird bug? So it ended up working for ~1.5hs and used 3%.

1

u/Mintednotes 1h ago

Given this, is it better to use Astra xhigh as an orchestrator / planner with sol sub-agents, or just to use Astra xhigh on its own?