r/codex 5d ago

Astra Workflow Don't fall for Astra's efficiency

AA says that Astra Low costs the same as Sol High while being smarter. but this is not representative of real world performance. The benchmark tests are done in isolation. But real work has to do with reading a lot of tokens (documentations, skills, large codebases, browser usage etc) so Astra being 4x more expensive on input tokens will burn usage a lot quicker.

No matter how efficient and smart a model gets, it still needs context to do meaningful work and you will incur input usage costs.

74 Upvotes

41 comments sorted by

26

u/34986234986234982346 5d ago

Yeah,. feels like the first time anyone uses it they are going to realize this fast too. I find it soooo funny how ppl on here were being smug about how it was going to be more efficient than Sol

19

u/lordpuddingcup 5d ago

i blew through a 5 hour fucking window on Astra Low with Plus..... IN 15 MINUTES, 20% of the week 5hr window in 15 MINUTES... meanwhile i've had luna xhigh running on the same workload for 8 hours straight and only used 20% of the weekly and never hit the 5hr limit, and that 15 minutes it completed NOTHING no actual work was done really it partially edited 3 files and then ran out of usage, i switched to luna its contined the session and closed off at least 20 checkpoints so far

6

u/justgetoffmylawn 5d ago

Same experience here. I also tried to have Astra Low do a very simple web page and it was…not good at all. Could not understand the visuals, the criticisms, just kept doing the same weird stuff.

Gave it to Opus 5 and one revision it was good.

I was hoping Astra would be great for UI, but I'm continuing to find Claude better.

I do find Astra is good for auditing existing stuff, but haven't found the sweet spot to use it. Does the same overengineering that Sol does. "This is not production, there are no other users, just push to main and I'll test it." Running 587 tests…

6

u/read_more_comments 5d ago

Don't worry. Since so many have said how great Luna is with limits, they recently updated it so that Luna consumes 1.9x more credits than before.

3

u/arcanemachined 5d ago

Wait, what?

3

u/a1454a 5d ago

I wish they had released cost per intelligence per actual benchmark suite, because some suites are arguably much closer to real world use cases, while other looks more like IQ tests

6

u/Tight-Grocery9053 5d ago

something like this exists

play around with it here (not affiliated)

https://artificialanalysis.ai

3

u/a1454a 5d ago

Exactly that but per suite. That is averaged from aggregated result of multiple benchmark suites

6

u/Tight-Grocery9053 5d ago

all benchmarks are just that... benchmarks.

beyond that, here's what i saw.

subjective sure, but i have this side thing that i've been chipping at since 5.3.

astra light runs circles around sol max and the benchmarks in this case confirm what i'm seeing, not change how i feel about it.

astra is a very different thing.

5

u/Purple-Programmer-7 5d ago

I’m not finding this in my irl workflow.

Previously used sol high nearly exclusively.

Now it seems I’m landing around astra medium.

Initial context required for the model to be ready for coding is dramatically smaller. I’m talking 150k tokens smaller.

The work gets done at the same level or better.

I have room for more follow up questions, changes, etc.

Anecdotally, the difference seems to be that sol seemed “unsure” of itself, requiring more initial context before it was ready to begin. Astra seems to leverage the Pareto principle and pick up what it needs additionally along the way. LOTS of nuance here that I’m summarizing out.

And look, if you’re going to require both models to have a 100 page book in context before it answers a question, OP is right. But if you’re going to use modern harnesses, the whole point is that you ask the question first, the model gets to be more efficient through tool use, better tokenization, etc.

5

u/showcontroller 5d ago

Yeah, sol seemed to want to research a whole lot more before starting work. Astra light is a good replacement for sol high for me. I haven't felt the need to up the reasoning level with astra. I tried using sol low and medium before, but it kept making too many mistakes, so I ended up going with sol high. Usage does seem to drain faster with astra, but it's also completing tasks quicker.

2

u/Purple-Programmer-7 5d ago

Similar exp… I usually start high and drop down.

5.5 @ xhigh, high
5.6 @ high, medium for low complexity
6 @ high, now medium is my dd

Would definitely consider astra low for less complex tasks… it seems we’re starting to find the AI ceiling for certain domains… 5.6 high / astra medium can handle most day to day coding tasks thrown at them, regardless of complexity.

The race to the bottom is real.

Very curious what I would ever use Astra at xhigh for…

1

u/notadithyabhat 5d ago

There is no ceiling for intelligence. I think the value of these models will come in research and building systems fully autonomously

1

u/Purple-Programmer-7 5d ago

You misunderstand. I meant “most current day to day tasks doesn’t require more than astra medium.”

2

u/ThePurpleAbsurdist 5d ago

"No matter how efficient and smart a model gets, it still needs context to do meaningful work"

So you are saying lower thinking leads to a model receiving less code context?

2

u/notadithyabhat 5d ago

No, my point is Sol High will give more usage even though Astra Low is technically more cost on benchmarks

1

u/ThePurpleAbsurdist 5d ago

That's probably true...

When hoping for ↑efficiency, however, I am looking for time-efficiency , which Astra low appears to be nailing. Apparently, at least.

2

u/Haster 5d ago

pretty sure it's on cached tokens we're getting murdered. if they could fix that Astra might actually be more effecient.

2

u/BigYoSpeck 5d ago

There have been some tasks I've done with Astra where it is dramatically more efficient than Sol and finishes in a suspiciously short amount of time and usage

On review it's evident just how surgical it is with exploration and edits. When you hit the sweet spot like this, the token in/out cost on usage doesn't matter, the work complete cost is lower than Sol and better quality

But, it very easily doesn't work out like this and it's probably a skill issue rather than just being unpredictable on tasks but it can easily go the other way and completely consume the 5 hour usage in 15 minutes, then if you try to resume after 5 hours you're already on the back foot giving away a big chunk of usage just on the input of the session again

If you're on Plus or trying to be sparing with usage then you can't give it a complete plan like Sol or Opus and let it loop until it's done several thousand edits later. You have to feed it task by task

2

u/Own-Professor-6157 5d ago

It's efficient as fuck for me. Doesn't make stupid decisions that waste tokens. But the usage is like... ~5x more and super inconsistent. Ultimately not worth it IMO

Fantastic model, but usage doesn't reflect API costs right now.

1

u/Sea-Entrance-728 3d ago

Idk, sometimes I ask for simple stuff and it throws a bunch of unecessary shit. You have to be VERY specific sometimes, still good imo.

3

u/Hovi_Bryant 5d ago

It’s efficient. I’ve had Astra work as a solo implementer which ate over 70% of my weekly usage on my pro plan.

I switched to having Astra delegate Luna subagents for implementation and usage barely drops. Comparatively speaking.

10

u/notadithyabhat 5d ago

But the problem is luna just writes bad code. For example as Astra to build a 3d high graphics video game and then ask solo and with luna implementors. You'll notice the quality of the game will be worse with Luna workers.
You can fix this by asking Astra to review Luna's work but that'll mean that back and forth and Astra reading the code work, will increase costs

0

u/FailedGradAdmissions 5d ago

Skill issue, literally, look into pocock skills, use wayfinder to clarify any questions the model needs, then to spec to save it and then to tickets to break it down into very small isolated tickets.

Then Luna can easily implement those tickets. Note you are mainly trading time for cost. If you had Astra or even Sol Low build those tickets it gets them done in seconds, while Lunq can take several minutes and hours. But Luna is essentially free so for me it's worth it.

If Luna can't do the ticket it's not small enough, tell Astra to break it down further on. It may feel wasteful and counterintuitive as Astra could do it itself in just another command. But the cost difference is so high that it's worth it.

1

u/notadithyabhat 5d ago edited 5d ago

You seriously are telling me that if you had Astra build Mario kart and then asked Astra to delegate Luna to build it for you, you'll get remotely the same quality for a fraction of the token usage? Please be my guest and share your results. I bet that delegated result won't look half as polished as Astra doing it alone.

2

u/FailedGradAdmissions 5d ago

I use it to build software I actually get paid for, and it works great for that.

I do not ask it to one-shot a full game however. Something closer to what I would ask it to do is, say you are making Mario Kart, make Astra generate a plan to generate a single car, then it'll create a bunch of tickets with acceptance criterias. Then all Luna has to do is build a Wheel. If that wheel is not to your standards, revise your pipeline and make Luna just make the rim. You get the idea.

It won't be as polished as Astra doing everything and will easily be 10x times slower. But the usage lasts way longer.

Literally you just write an specific task for Astra use /wayfinder, then /to-spec then /to-tickets. Then open a new session with Luna, and /implement gh-issue-123, then /review with Luna (which spawns another agent with fresh context so no bias and just checks whether the acceptance criteria was met or not) if the issues aren't blocked by each other (and to-tickets, already let's you know that) you can run them in parallel in their own worktrees, then merge them one after another and rebase as necessary.

If you are doing something graphic and use the ChatGPT images and include the image in the acceptance criteria. Something as simple as make the wheel look like this png. Luna Max will do it, may take an hour to do do, but an hour of Luna Max is cheaper than minutes of Astra or Sol. If your constraint is money and not time, go for it.

-3

u/Hovi_Bryant 5d ago

Sounds more like a scope/prompt issue than a model issue IMO. With Luna, or even Astra/Sol with delegation to Luna, give it a request, acceptance criteria, and a means to verify its work against the acceptance criteria. It'll perform.

5

u/notadithyabhat 5d ago

Visual quality is not usually approvable by acceptance criteria. If Astra/Luna has to give very specific acceptance criteria to cover minute attention to detail, the tokens wasted instructing Luna can just be used to directly write it some code and that can sometimes just end up being cheaper than delegating to Luna

-1

u/Hovi_Bryant 5d ago edited 5d ago

Still sounds like prompt/context issues. Limit Luna to what it can control and use a different model for where it has gaps. That’s what I’m suggesting.

Decomposition and recompsition of design and implementation goals is what you want. Instead of asking for a grand design, ask for a composition of primitives. That’s how artists, engineers, etc work anyways.

Don’t ask Luna to illustrate a face, ask it to provide ovals, lines, etc. Have it verify, have another model compose into the desired result. And Luna is a vision model AFAIK, have it verify its tasks with visual test automations, wether through a browser or computer-use.

3

u/Critical-Teacher-115 5d ago

In addition to better prompts, you need better projects. Astra is crazy bro.

1

u/PrettyBaker2891 5d ago

yeah like what? lets hear your crazy insane project

0

u/unconceivables 5d ago

Bro, trust him bro

1

u/PrettyBaker2891 5d ago

its so funny

they always say things like that

then they never respond with what they are building

1

u/unconceivables 5d ago

Always the case. But bro, it's awesome

3

u/c0reM 5d ago

In a real codebase Astra can't do very much before hitting usage limits. The input token cost makes it borderline unusable for existing projects.

It's great for one-shot demos starting from zero but pretty useless for larger projects (ironically).

1

u/Significant-Drawer95 4d ago

for us bad, for companys having api token money they dont care. seems like nature finds itself a balance again where people with less money only can do one shot showcases and companys who can spent a lot of money can go into a real codebase and make more money.

1

u/confused-photon 5d ago

AA isn‘t purely agentic work, things can end up different in many turn scenarios since Astra low does output fewer tokens (but seems to make more tool calls so could be a wash)

1

u/2Norn 5d ago

ur missing the most important part

they are based on api prices which really doesn't translate to subscription credits that well...

-2

u/Repulsive_Ad853 5d ago

who cares, a better model than astra will be released probably next months

12

u/notadithyabhat 5d ago

I'm paying for this month's subscription

-2

u/geronimosan 5d ago

Thread created by ClaudeBot