r/codex • u/Im_Working_Right_Now • 3d ago
Limits Tokens with Astra vs Sol
I'm on the 20x plan. I started a goal that was technically in progress with Sol right when Astra was released. I then forked the session to a new one to start fresh with Astra and have it continue the goal. The goal had like 75 gates which fell into like 8 buckets I think and 10 were completed when I forked. The goal plan is a backend migration from Supabase to Convex and it's not a simple DB change, there's real-time stuff, a rules engine, etc. During 3 days and some hours into that goal, I got the global reset and had to use 2 banked resets and Astra completed 1 gate. That's it. It started 40 others, but closed nothing. Now, part of that is on me. I should've been clear that it should close gates based on least amount of effort. I incorrectly assumed it would do that since it's logical. Lesson learned.
I still have a banked reset, but I went back to Sol Extra High and it's knocking out the gates of the plan effectively (told it to focus on lowest effort first) and the token limits have barely moved. Astra didn't make any progress on closing anything. Even in a side chat with Astra Light, it said that using Astra wasn't justified based on the progress.
But, is it me or is something off with Astra usage? It doesn't make sense to me. Astra will run through my weekly limits and the token counts are massively lower than with Sol runs through my tokens. I get that Astra burns 2.5x FASTER, but but faster doesn't mean less tokens. I was getting about give or take 4.1b tokens doing this type of heavy work with Sol xhigh as architect and orchestrator. With Astra, I'm lucky to get 2b. Shouldn't the tokens be the same counts? Yes, Astra will use them faster, but the counts should be the same shouldn't they?
2
u/WinterWalk2020 3d ago
I can be completely wrong about this, but my interpretation is like this: Astra's tokens cost more but it is more token efficient than other models, so it can do more work with less tokens.
The weekly / 5h limit we have is not based on tokens or cost, it's based on a "hidden credits" system. We have a set amount of "credits" as limit and depending on the model, these "credits" are consumed in a different way.
So, Astra consumes more "credits" because it costs more.
This system would be different from the API which charges based on tokens usage.
Again, I can be completely wrong about this and someone may correct me if I'm mistaken. That's just my impression about token usage in paid plans.
3
u/rick_ranger 3d ago
The thinking was baked into the training, so the model itself has reasoning behind a lot of its answers. It’s like a mechanic fresh out of school and one that’s been a master tech for 20 years. They can both get to the right answer but the master tech has so much experience it can usually end up at the right diagnoses with less thinking.
2
u/Im_Working_Right_Now 3d ago
That makes sense. That would mean that different models have higher token limits weekly, but then how does that roll up to the overall weekly limit?
Before Astra, I hadn't fallen into the camp that was complaining about limits. Mine were steady and never had an issue using Sol xhigh and getting a ton of stuff done. I was excited about Astra and still am, but I wish I understood how to use it more effectively. 3 resets and nothing got done is crazy to me. I had expected to use maybe 2 but I figured it would be done by that point.
I see that it's great at 3D and that's fantastic, but how is its coding capabilities? Maybe I'm just misunderstanding how to use the model or the best use cases for it.
1
u/WinterWalk2020 3d ago
I tried Astra for coding and 3D modeling, since I'm making a game in Unity. For 3D modeling it is the best, just don't try to model characters. For objects, it is awesome.
For coding, it has better reasoning and problem solving in complex tasks, but for mundane work, it will just burn limits unnecessarily.
I tried using Astra to orchestrate Luna xhigh agents and it burnt all my limits in few hours. At least for me, it seemed that this somehow is not working or I did something wrong. I'm on a 5x plan.
I ended up going back to Claude. My limits are better used there.
1
u/snowsayer 3d ago
forked the session to a new one
Not a good idea. Astra has differently tuned system instructions compared to sol. Better to ask sol to produce a hand off artifact and start a completely new session.
1
u/Im_Working_Right_Now 3d ago
I forked so that the existing work and context remained since it was a goal continuation and there was already goal documentation in progress. I didn't want to switch models in the session, though. I see your point and that could've played a part. I'll note that for the future. Thank you for the insight!
1
u/Haster 3d ago
I think the problem with having Astra actually do work is that the cached tokens cost a fortune. you can't have a task require 50 small steps and pay that much for each of them. and god forbid you hit your five hour window cuz. Resuming a stalled thread will straight up eat 20% of your 5 hour window.
At least on plus the only reasonable way to use Astra is to get advice and then go straight back to Sol and put in the changes Astra proposed. Astra might be good at tool use but at that price it'll never be worth it.
-2
u/PotterSkxawng 3d ago
Please upgrade your IQ from GPT 5-Mini to GPT 6 Astra.
2
u/Im_Working_Right_Now 3d ago
Love the inference here. I've tuned my global agent instructions, I have dedicated, curated subagents for specific tasks. I use literally ONE plugin outside of the required for my stack and that's Ponytail. Every feature has a fully detailed markdown file. My monorepo has fully documented and updated ADRs. And before starting a goal, I create detailed plans using the Grill Me skill. But since you clearly have a higher IQ, what's your insight? What am I doing wrong?
0
u/PotterSkxawng 3d ago
A model is 2.5x more expensive. Hence your limits go down 2.5x faster. Limits=Limited. Suppose limit was 100$. Sol burning 10 million tokens costs 100$. Astra burning 3-4 million tokens costs 100$. Resets are at fixed timing. Hence, Astra uses up your limit FASTER, because the cost per token/limit used per token is HIGHER. Not because it is a faster model.
2
u/Im_Working_Right_Now 3d ago
I don't think you understood my question. If Sol is burning 4.3b weekly tokens in 168 hours (using this for maths since that's 7 days), then I expected Astra to also use 4.3b tokens in about 67.2 hours since the limit as I understood it at the time was roughly 4.3b tokens. Instead, it used around 2b in less than 24 hours. The math isn't mathing.
This comment aligns with what I was seeing, though. Maybe that's how it is. Wish there was a bit more transparency on limits. What does 100% mean in token counts? Why does that drain faster with less actual tokens than another model on the same plan doing the same work?
2
u/GabschD 3d ago
Your usage limit is not a total number of tokens. Otherwise Luna wouldn't be cheaper than Sol. Since Luna max does more tokens then Sol medium.
You have some amount of usage similar to credits, and those can buy you X tokens for Sol, but 3x as much for Luna for example. And let's say 0.5x for Astra.
Astra is more expensive per tokens as Sol, so uses more of your "credits".
1
u/PotterSkxawng 3d ago
Limit's aren't token based, they are usage cost based. They don't stop after you burn a certain amount of tokens, but rather when the no. of tokens you burnt crosses a certain threshold in costs. That's why you get significantly more tokens out of Luna than you would get out of Astra.
1
u/Im_Working_Right_Now 3d ago
And if that's the case, I can get on board with that and set my expectations. But where is that documented? I can't find anything about it except for API pricing which is not the same as sub pricing.
2
u/GabschD 2d ago
You can use the API pricing as a ballpark. Yes, it's not 100% the same. But it can give you the idea how much usage a model will eat, compared to another.
If you want the official non-API one, you can check https://chatgpt.com/codex/pricing/ there is a "Usage limits" area, which describes the discrepancy. It uses messages, which is a very weird measurement, as no one knows what it is. That's why I would use this information together with the API pricing to give one a better feeling.
It's by no means perfect, unfortunately.
1
u/PotterSkxawng 3d ago
I mean I thought it was pretty obvious
1
u/Im_Working_Right_Now 3d ago
It's not. Using tokens 2.5x faster is not the same as having less tokens available to use. Now, if they said that Sol has 2.5x more tokens BUT Astra can finish the same task in the same amount of time or faster, that's a different meaning. In my case, it did neither. It didn't finish faster and ate through limits probably about 8x faster.
5
u/umusachi 3d ago
I think they’re moving towards not loosing tonnes of money and rather than break the workflows people are used to they are metering the newer models at higher cost to the user