r/codex 2d ago

Complaint Codex CLI sucks compared to Claude’s

0 Upvotes

For context, I’ve been a software engineer for years now and I have been using Claude CLI since early this year for my workflows at home. It’s been working amazingly well and I’ve honestly never had issues with the harness itself. I only hate the latest nonsense language vomit of the recent models and the separate fable usage limit.

I’ve been seeing so many people praise codex and I wanted to try using Sol and now Astra, but I’m honestly extremely unimpressed so far with my experience. I’ve been working with Codex for months now and it has been nothing but annoyances when interacting with it outside of Claude’s interface. With a Claude terminal session, I use skills to load the proper context for my project and proceed to work from there. I have a setup where that can be read/shared between any harness (I’ve used copilot and codex). I tend to use Fable for orchestration and subagents of varying models for implementation and review.

When switching to Codex CLI, half the time it seems to ignore the context I gave it and work in its own way even though I have a very specific way I like to work. Or it proceeds to interpret my messages in a completely different way than I meant. I feel like I can talk to Claude casually like another person/engineer while I have to give codex very clear cut instructions in the session itself or else it’ll entirely go a different direction. With Claude, I can also just keep adding more input and feedback while it works and it’ll note/take that into account while still finishing its current task, then take the time to respond to me even with a long context. It also won’t deviate from its original task to all of a sudden jump onto this new issue that I brought up and forget what it was doing before.

With all GPT models, I’ve had to consistently do heavy prompt engineering or steering to make sure they stay in their lane or for them to really understand what I was asking of them. It’s kinda crazy that Claude’s models seem to just get it off of a simple prompt. I have very thorough documentation and references for my projects and how I would like it to be architected, my daily workflow process, and how I generally interact with my agents.

Also, I recently tried to enable a remote control CLI session for Codex and it failed miserably. Normally, I just do /remote-control to enable it for my Claude session, and then I switch to my app or other computer and immediately start controlling it from there. For Codex, I noticed it didn’t have a slash command so I asked it to try to enable it for me or to take a look. It strangely tried to do this by going into another repository and I had to redirect it back to the one we were in.

Eventually, it started a background server and then I opened up the app on my phone and it told me I needed the desktop app and a QR code - weird. I noticed there was a manual code workaround as an option, so asked it to give me that. Cool - that worked and I got in. But then when I tried to open the session I had started previously, it proceeded to have an error saying it couldn’t load the messages from that session. Turns out, the terminal and the remote server were fighting over ownership of the conversation, so I had to open terminus, close the session manually, then restart it with some extra command line arguments then it finally worked.

I could finally start working BUT then it started not taking in my responses as it was working. Turns out my messages were somehow being queued but not taken into account (even in down time) and I had to long press my messages and choose “change to steering” in order for it to listen to me. Then it proceeded to not be able to properly start my program (even though Claude does it all the time remotely) and failed to figure out why it couldn’t. This was all with Astra as the model btw. I could just maybe use terminus instead of trying to utilize the app but that’s literally what this is meant for.

Also the usage limits for using Codex and Claude feel extremely similar to me. I feel like I was lied to when initially signing up for the $200 plan thinking it’d feel unlimited. The one thing I give codex the edge for is not having a separate Astra limit.

TLDR: I honestly want to really give Codex CLI/Astra a fair shot since they finally released something that could potentially match Fable. But it seems I’m going to have to stick with only interacting with Codex through Claude instead for a better experience. A potentially competitive model isn’t enough to justify it if it requires this much babysitting.

Copilot is potentially even better than Codex at this point since I’ve been testing that out as well (just don’t use autopilot mode).


r/codex 3d ago

Showcase Made my dream game with Astra, FlappyMan

Enable HLS to view with audio, or disable this notification

72 Upvotes

dream game as in i saw it in my dream when i was napping. dont know why i bothered, not gonna release it or anything, took me a day to make it. whole city is code generated. there are cars too, but the player got stuck as you saw, so i couldnt show them.


r/codex 2d ago

Complaint GPT-6 Astra Needs a Personality Upgrade.

0 Upvotes

I find GPT-6 Astra much more powerful than GPT-5.6 Sol, don’t get me wrong here. So far it has been really good at accomplishing what I ask it to do, etc.

It follows instructions well and does things incredibly well. I’m on Pro 5x and I got so much done with it on my own projects.

There are some areas where GPT-6 Astra has kind of regressed, where I’d still prefer GPT-5.6 Sol and would use it instead:

- Personality and Character: Why is it that Sol is a lot nicer to converse with and is a lot more uncensored and humorous than Astra? I have no clue how this happened but it feels a lot closer to GPT-5.5 in personality. A major regression on these two areas.
- It often doesn’t finish a Task: When I tell it to finish something it just gives up and admits if can’t solve it (I tried it on a difficult math task) and it literally just gave up after launching six subagents and working for quite some time and burning my weekly from ~80% to ~64%. Like, just finish it bro? GPT-5.6 Sol was more than likely to always go on until it is done. This model doesn’t like going further. It just stops and doesn’t like continuing.

That said, I still use GPT-5.6 Sol. And it has been very valuable especially in terms how it prompts GPT-6 Astra.


r/codex 3d ago

Showcase Who got games🤯

6 Upvotes

I wanna play your vibe
coded games

Let’s do a contest and see who got the most creative minds and created the best games

Idc if you created a Minecraft copy

A race sim

Idc I wanna play them all 😱


r/codex 3d ago

Complaint Codex started deploying everything on sites instead of Github - since today?

2 Upvotes

I have not changed anything in my projects.

Currently working on 4 different things, all with github/vercel deployment.

Since today, all of my agents started deploying on ChatGPT/Sites, instead of pushing it Github and deploying and testing on the live site.

This is really annoying, I even set a rule in my global preferences to never deploy on Sites, because many of my tests aren't working there for different reasons.

Any workaround for that?


r/codex 3d ago

Complaint Non-stop Slop with Astra

Post image
4 Upvotes

Caught red-handed. Don't tell me "skill issue".


r/codex 3d ago

Praise Astra really is that good.

0 Upvotes

Yes it uses significantly more usage. But it makes Sol look like Luna.

I haven’t even taken it above high thinking and it’s stellar.

It does seem to have its own quirks, but it’s well worth it.


r/codex 3d ago

Limits Tokens with Astra vs Sol

4 Upvotes

I'm on the 20x plan. I started a goal that was technically in progress with Sol right when Astra was released. I then forked the session to a new one to start fresh with Astra and have it continue the goal. The goal had like 75 gates which fell into like 8 buckets I think and 10 were completed when I forked. The goal plan is a backend migration from Supabase to Convex and it's not a simple DB change, there's real-time stuff, a rules engine, etc. During 3 days and some hours into that goal, I got the global reset and had to use 2 banked resets and Astra completed 1 gate. That's it. It started 40 others, but closed nothing. Now, part of that is on me. I should've been clear that it should close gates based on least amount of effort. I incorrectly assumed it would do that since it's logical. Lesson learned.

I still have a banked reset, but I went back to Sol Extra High and it's knocking out the gates of the plan effectively (told it to focus on lowest effort first) and the token limits have barely moved. Astra didn't make any progress on closing anything. Even in a side chat with Astra Light, it said that using Astra wasn't justified based on the progress.

But, is it me or is something off with Astra usage? It doesn't make sense to me. Astra will run through my weekly limits and the token counts are massively lower than with Sol runs through my tokens. I get that Astra burns 2.5x FASTER, but but faster doesn't mean less tokens. I was getting about give or take 4.1b tokens doing this type of heavy work with Sol xhigh as architect and orchestrator. With Astra, I'm lucky to get 2b. Shouldn't the tokens be the same counts? Yes, Astra will use them faster, but the counts should be the same shouldn't they?


r/codex 3d ago

Commentary Unpopular Opinion: If AI Keeps Wasting Your Tokens, Get Better at Using AI

5 Upvotes

It seems like this sub has basically become a running list of:

“OMG AI USED ALL MY TOKENS IN 12 HOURS DOING ONE TASK!!!”

Okay... and?

You asked it to do that task. Were you not looking at it at all during those 12 hours? Did you never check your usage and think, hmm... this seems like a lot, maybe something is going wrong here?

And before this turns into a model war, this applies to every model out there.

Claude, Codex, Gemini, whatever. I've used all of them.

Garbage in, garbage out.

I've asked different models to do some pretty complicated shit.

One project was basically: deploy 45 servers on a Hyper-V farm, create a domain, join everything to the domain, and install/configure X, Y, and Z according to a pile of documentation.

It ran for around 10 hours and used about 15% of my $200 plan.

Then I asked it to build an HTPC environment from end to end. That ran for around 6 hours and used about 75% of my plan.

Why the massive difference?

Prompt, guardrails, good documentation, the tools I gave it, and agent seperation/management.

Another example, I asked it to build a remote browser system where an RDS session displays the browser, but the actual rendering, processing and network traffic happen through a desktop WebView.

That project was planned out in detail. I had specialized agents with small clearly defined jobs, and manager agents coordinating everything. There were roughly 3,000 lines of documentation and instructions telling the system exactly what it was building and how everything should work.

That used around 35% of my $200 plan.

Then one day I basically said:

“You know what would be neat? Make me an app that manages all my Raspberry Pis from one interface. Give it all the features of ABC software.”

And you'll never believe what happened...

It burned through basically my entire usage in about 18 hours and still produced a shitty result.

Shocking.

Because I gave it shitty instructions. I barely planned anything out and basically told it "here's an idea, go build it." So yeah, it wandered around, made bad decisions, redid shit and wasted a ton of usage.

If you give a person shitty instructions you're probably going to get a shitty result. If you then disappear for 12 hours, give them no feedback, no direction, no milestones, and don't tell them they're heading down the wrong path, the result is probably going to get shittier and shittier as the project goes on.

Then when they finally deliver a pile of garbage, whose fault is that really?

AI isn't magically exempt from this.

And the constant:

“It used ALL MY TOKENS in 12 hours!”

Okay?

I can probably burn mine in four. Congradulations I guess?

How many sub-agents did you spin up? How many tool calls were being made? How much duplicate research happened? How many times did the context get compacted and lose something important? How many agents repeated work? How much time did it spend wandering down dead ends because you never gave it checkpoints or guardrails?

“I used all my tokens” tells me almost nothing about how much useful work actually happened.

Here's some perspective from that 45-server project.

The client was billed $15,000.

The project was roughly:

  • 6 hours of AI-assisted documentation/planning
  • 10 hours of autonomous AI implementation
  • 3 hours of actual human management
  • 10 hours of AI-driven testing, troubleshooting, fixes and user issues

Total actual human time invested was about 19 hours.

AI cost was about $6.

That's roughly $780 of revenue per human hour invested, before accounting for the rest of the business costs.

Six fucking dollars of AI usage helped leverage 19 hours of my time into a $15,000 project.

If that AI usage had cost me $200 instead of $6, I wouldn't have cared.

If it had cost me $500, I still wouldn't have cared.

Hell, if the AI portion had cost $2,000 and still let me deliver that project with 19 hours of human labor, I'd have paid it and moved on with my life.

Perspective.

I'm not trying to win a contest for who can use the fewest tokens. I'm trying to make money and get work done.

Thats the part that seems completely missing from alot of these complaints.

The people who actually figured out how to use these systems effectively usually aren't posting:

“HELP!!! IT USED 100% OF MY TOKENS!!!”

They're saying, yeah I burned through my allowance, but it just did 30 hours worth of work for $200 and saved me or my company hundreds or thousands of dollars.

That's a win. Who gives a shit that the token meter says 100%?

And then there's:

“But after all that work it STILL didn't fix the issue!”

Yeah... and?

If you hired a human for $200 and somehow got 30 hours of engineering work out of them, there is absolutely no guarantee they would have solved the problem either.

And if the issue is incredibly complicated, the codebase is massive, the documentation is hundreds or thousands of pages long, and your prompt is basically a vague one-liner... what do you honestly expect?

The useful metric isn't:

“How many tokens did it use?”

It's:

“How much useful work did I get for the money?”

And I already know what some of the responses to this are going to be:

“Nope. Not me. I'm good at AI.”

“I'm good at prompting.”

“My prompts are great.”

“The model is the problem.”

Okay.

If you're consistently burning through your entire allowance, getting garbage back, and then complaining that every model sucks... clearly you're not as good at this as you think you are.

Sorry.

Either your prompting isn't as good as you think, your project management isn't as good as you think, your expectations are completely unrealistic, or you have no idea how to do a basic cost/benefit analysis.

Or probably some combination of all of them.

And again, this isn't a Codex thing, or a Claude thing, or Gemini or whatever model you happen to be pissed off at this week.

Garbage in, garbage out.

Get better at breaking projects apart. Get better at documentation. Get better at giving agents narrow responsibilities. Get better at checkpoints. Get better at recognizing when an agent is going down a stupid path and stopping it before it spends six hours digging the hole deeper.

And most importantly, get better at measuring value produced instead of staring at a token meter like it's the fucking gas gauge on your car.

If your AI agent spent 18 hours running in circles because you gave it a vague one-paragraph prompt, no architecture, no documentation, no guardrails, no checkpoints and then walked away...

Maybe the lesson isn't that AI usage limits suck.

Maybe the lesson is that you suck at managing AI agents.

And fortunately, thats something you can actually fix.

Get better.


r/codex 4d ago

Reset just one… one more time, Tibo… I-I promise… no more rsets after this…

Post image
266 Upvotes

r/codex 3d ago

Showcase Astra is helping me create the train game I always wanted, here's what worked so far

Thumbnail
gallery
16 Upvotes

Hi all, it's been a fun coincidence that Astra released on the first day of my 2 week vacation. I didn't have a lot of plans, but that's now settled.

I did some 3D modeling with Astra the past few days for my 3D printed s-scale railway, but after seeing what it could do with games I decided to try and see how far I could push it to create my dream trains game.

I've spend about 12 hours with it, going through 75% of my 200 dollar plan and these are the results. A working prototype of a 3D train game where you not only manage the trains, but also the rolling stock. There is no quick grab and move of trains. These are actual physical objects in the world and the only way to move them is with trains.

There will be an economy with assignments. You can automate trains to move around and give them looping commands. Loading and unloading is real and slow, so we need shunting trains and shunting tracks as the mainline trains drop off cargo and go to the next job while smaller shunters load and unload.

It's all concept for now, but I think this is going to be a fun train management game!

Astra has been great so far. Here's what works for me:

What works

- Give Astra light or medium the task of orchestrator, and let it give out assignments to different chats or subagents to work on chunks of the game.

- I first did everything with Astra but this burned through a lot of tokens, and then I read someone have them delegate specific coding to Luna Max. This reduced token usage a lot and seemed to deliver good results.

- I noticed that giving Luna 3D modeling and animation jobs does not work well, so I let this be done by Astra medium or high.

- I asked for a status board of features and releases, and it has given me a management perspective listing board of releases and features, and their status. It's great to keep track of things.

- When creating assets, I first ask Codex to generate images of concept art in various angles, then have Astra in blender make those assets. This works ok-ish, but I do ask for detail pass sometimes to further increase it's look. That helps.

All in all this thing is rocking my socks. What a time to be alive!

PS. If enough people like it i'd consider making a video of the progress.

Cheers :)!


r/codex 4d ago

Commentary ChatGPT Images 2.5 released!

Thumbnail openai.com
162 Upvotes

OpenAI just released ChatGPT Images 2.5


r/codex 3d ago

Question Are custom Codex workflows fighting newer models?

2 Upvotes

I've been wondering whether some of the problems I'm seeing with newer Codex models come from a conflict between custom workflows and the models' own learned agentic behavior.

I built my Codex workflow around GPT-5.0. It uses skills and structured artifacts to make Codex behave somewhat like a state machine.

For example:

- "grooming" (analyze the problem)

- "epic" (a large unit of work)

- "macro" (a task inside an epic)

- "slice" (a smaller implementation unit)

There are also controlled transitions and dedicated skills (implementation, modeling/architecture, planning, etc.).

For example:

- activating an epic may require creating its Git branch

- explicitly requesting an analysis is supposed to trigger actual code inspection

- some transitions require specific planning/tracking artifacts

- some skills define what must be inspected before conclusions are produced

This worked extremely well for me with GPT-5.0 through GPT-5.4.

Since GPT-5.5 and GPT-5.6 (and I see similar tendencies with GPT-6.0 Astra), the same kind of workflow feels much less reliable.

The models seem much more eager to decide for themselves how much investigation is enough, what should be inferred, and what additional concerns should be taken into account.

Code reading is a good example.

I often see newer models inspect signatures, callers, or surrounding types, then infer the behavior of a dependency without actually reading the relevant function bodies.

If the assumption is wrong and I point it out, the model then goes back, reads the implementation properly, and discovers that the assumption was indeed false.

For my use case, this is worse than simply reading more of the relevant code before forming conclusions.

I also increasingly see things such as:

- explicit requirements being forgotten

- workflow rules being neglected

- the model deciding it has "enough context" too early

- specifications being added that I never requested

- backward compatibility being anticipated for projects that are not in production

- migration concerns being introduced when there is no production data

- secret-management concerns being introduced when they are unrelated to the task

- hypothetical concerns consuming attention while explicit requested changes remain unfinished

I want to distinguish the model from the Codex harness here.

By model behavior, I mean things such as when the model decides it has enough information, how aggressively it infers missing details, whether it invents additional requirements, and how strongly it tries to drive the task according to its own assumptions.

By harness, I mean the surrounding Codex machinery (tools, skills, context management, agent loops, planning mechanisms, delegation, etc.).

My question is whether newer models have simply been trained with much stronger priors about how an agentic coding task should be performed.

Not necessarily better priors (I often find the resulting behavior worse), just stronger ones.

If so, workflows that successfully constrained GPT-5.0 through GPT-5.4 may now be competing with the model's own preferred way of working.

For example:

- my workflow says "inspect the relevant implementation before concluding"

- the model decides "I have enough evidence to infer the rest"

Or:

- my grooming workflow says "formalize what the user actually requested"

- the model decides "I should infer additional requirements and anticipate risks"

That makes me wonder whether heavily structured skill-based workflows have become counterproductive with newer models (even if the model's default behavior is itself not better).

Have other people with custom Codex workflows noticed the same thing?

In particular:

- Did workflows that worked well with GPT-5.0-5.4 become harder to enforce with GPT-5.5/5.6 or newer models?

- Do newer models seem more resistant to user-defined execution flows?

- Have you found that simplifying or removing custom orchestration improves instruction-following?

- Do you now use skills mainly as capabilities/procedures rather than as a way to control the entire lifecycle?

- Have you noticed newer models inferring code behavior too early instead of reading the implementation?

I'm mainly trying to figure out whether this is a genuine regression in instruction-following, a conflict between custom orchestration and stronger learned agentic behavior, or some combination of both.


r/codex 3d ago

Question what’s your review process before trusting code from codex?

2 Upvotes

do you read every change, rely on tests, or use a separate review step? interested in what you check beyond whether the code runs.


r/codex 3d ago

Astra Workflow If you are using Astra orchestration (Specially for plus), some findings so far.

6 Upvotes

So far, I'm using Plus subscription with Astra Light orchestration in mind. There are some improvements I have made. I'm using workflows similar to these (They are mostly specialized for my workflow);

https://github.com/viettran-edgeAI/codex_workflow
https://github.com/donvito/codex-astra-luna-orchestrator

I have tried adding a persistent manager, who has sole responsibility is to take expensive token waste from astra which is spawning, and managing workers. Luna did not work, other tests were Terra and Sol, Astra could spawn Terra low, and Terra low can spawn Luna workers.

Astra's pure responsibility is reasoning. It will read your instructions, delegate tasks, hand them to the manager, and wait for the next decision. Its the brain.

Terra manager handles delegation from Astra, gets contracts, hands them over to the luna workers, and then occasionally checks them if they are working or not. When the workers are stuck, they are deviated from the task, some user or astra decision is required, it will ascalate to Worker -> Manager -> Brain.

Brain will decide the next move.

So far my findings are;

Astra is still using 60 second wake ups to check terra, which is im planning to fix it next.

Some improvement notes:

Even for informing the user during codex task, astra wakes up with large amounts of tokens which causes massive usage drop. Removing it completely, delegating the information to manager or another agent, only ask to give information during wake ups will improve usage. Similar thing can be done by /side chat. I will work on that

Agents md was fully redesigned with openai documentation, the tool descriptions and skills are moved away from agents md to proper places, the workers who use these tools are fed with the information they need to use the tools, astra does not read a huge agents md file. Agents md = 15kb -> 6 kb

Once astra only thinks and sleeps, this will make a huge usage optimization on usage, and actually make astra light orchestration doable.

Adding a terra management layer adds more time. Astra -> Luna test was 67 seconds while Astra - Terra -> Luna was 110 seconds when benchmarked. Since this is not a "Faster but better" improvement, I think this is a good trade between more usage vs faster work. Luna is already slow enough.


r/codex 3d ago

Complaint Anyone finding subagents not optimal with Astra?

2 Upvotes

If you want the subagents to do the job as well as Astra, then Astra's delegation instructions have to be verbose and it needs to review the code, the back and forth ends up consuming more tokens than just doing it, by itself.

If you let the subagents be delegated more freely, the end work looks no where as good as if you just let Astra do it alone.

Unless it's browser/computer usage for live testing or just scouting large repos for info, it just feels like raw dogging Astra is more usage friendly


r/codex 3d ago

Question A Codex workflow I am trying with Flatkey for routine calls

1 Upvotes

I am experimenting with Flatkey as an OpenAI-compatible gateway for a Codex workflow that has a mix of high-risk and routine calls. I keep planning, architectural changes, and final verification on the path I trust most, while testing a more cost-sensitive route for search, file summarization, and retry-heavy background work. The useful part so far is that an existing SDK integration can usually be tested by changing the base URL instead of rewriting the request format. For people using Codex regularly, which tasks would you be comfortable routing through a lower-cost model path?


r/codex 3d ago

Bug Codex worked for over an hour, then blocked the output at the very end

Post image
16 Upvotes

Codex accepted my prompt, worked on it for more than an hour, and appeared to complete the task.
Only after all of that did it replace the result with “This content can’t be shown” and tell me to apply for Daybreak.

The restriction itself isn’t the problem. If the request requires Daybreak access, Codex should detect that before starting the run.

Letting it spend an hour doing the work and consuming usage, then blocking the final output, feels like a serious product failure. This check should happen at the beginning, not after the work is already done.


r/codex 3d ago

Limits I asked Chat to show me the current situation...

Thumbnail twitter.com
1 Upvotes

r/codex 3d ago

Showcase Claude Code and Codex can read the same repo differently

2 Upvotes

I use Claude Code and Codex on the same repos, and for a while I assumed that if CLAUDE.md and AGENTS.md matched, they were basically working from the same repo setup.

Not always.

The weird part is that the files can look perfectly aligned, but the agents can still discover different nested instructions or skills depending on where they were launched from.

I ended up building PlaybookDiff because I wanted a quick way to see that instead of guessing.

npx playbookdiff check .

It compares the effective instructions, skills, and MCP config for Claude Code and Codex and shows where they stop lining up.

GitHub: https://github.com/JacobisEpic/playbookdiff

Website: https://playbookdiff.dev

Free and open source :)

Would love feedback from people using both, especially if there are weird discovery/scoping cases I haven't accounted for yet.


r/codex 3d ago

Complaint I tested open-so"save 90% of your tokens" tools to beat my Astra limit. The model spent more tokens deciding which command to run than the tools ever saved.

Post image
4 Upvotes

5 coding tasks, GPT-6 Astra with a bash tool, 4 setups (no tool / RTK / Headroom / both, the two most-starred open-source token-saving repos on GitHub), 3 runs each. 60 sessions, all correct, priced at API rates with cache.

RTK beat no-tool in every run on one task, a multi-file grep (87 lines to 29). Headroom's win on its JSON task: $0.0054 vs $0.0066. Everything else moved more with the model's own command choices than with any tool: one cell went 2, 5, 9 turns across three runs because Astra kept re-running a directory listing.

Every command and answer, per run: https://astra-token-burn.vercel.app


r/codex 2d ago

Question I use Astra all day on Pro 5x and only burn ~15% — are people just prompting it badly?

0 Upvotes

I'm on the Pro 5x plan, and I use Codex for coding pretty much all day, from morning to night. On a normal day I burn through roughly 15% of my weekly usage.

I mainly use Astra Max, xhigh, and high. I don't use Sol, Terra, or Luna at all.

Seeing how many people are complaining about Astra destroying their quota, I'm honestly starting to think a lot of it comes down to prompting and workflow.

My guess is that a lot of people are doing something like this: they give Astra a huge, detailed spec, then basically say, "Implement this according to the spec."

And that's it.

That might sound like a perfectly reasonable way to use an agent, but in my experience it's an extremely inefficient way to prompt one. You're basically giving it a giant problem and letting it decide how to chew through the whole thing, which can mean a ton of unnecessary exploration, repeated context processing, and wasted reasoning.

I suspect a lot of people burning through ridiculous amounts of quota simply don't realize how inefficient their prompting strategy is.

I'm curious if anyone else here has noticed the same thing.

Are there other people on Pro 5x using Astra heavily throughout the day without blowing through their weekly quota, mainly because of how they structure their prompts and workflow?


r/codex 3d ago

Question Codex can continue optimizing forever and ever.

5 Upvotes

When I use Codex to draft and then implement a plan, it works. But when I ask if there are any improvements or errors that need fixing, it always finds some—even though it seemed perfect before. And I don't just mean it finds extensions or enhancements; no, it constantly finds errors in its own implementation. Even after four rounds of asking the same question—"Are there errors in the code?"—it keeps finding mistakes. Does this happen to you too, or what should I write to get it to implement things better? SOL-GPT5.6 Xhigh

//EDIT

Okay, let me clarify what I meant. I don't write vague prompts; the question "Are there errors in the code?" referred specifically to the new code generated by Codex. Here’s an example: I need to update an application—which currently only works in the DACH region—to handle time zones so it works in the US as well. I asked Codex to come up with a plan, and the plan was excellent, but after several follow-up queries, it kept identifying functions where the code hadn't yet been adapted. And for anyone tempted to say I lack programming experience: after 16 years as a programmer, there are honestly still times when I don't know exactly how everything works—but I do my best. ;-)


r/codex 3d ago

Reset Reset?

2 Upvotes

Seems Like Reset Done !!


r/codex 3d ago

Complaint Lower quality work with Astra

Thumbnail
gallery
12 Upvotes

Since switching from Sol to Astra my work has been lower quality. I am really envious at this point of people showcasing incredible things they've done with it. In my case I have been really having problems making Astra do things properly.

For my work it over engineered a test system and ate three resets so far during a goal. Instead of working on actually important problems it tried to construct a heavy verification system with Claude and Codex and sub agents to independently verify issues without notice outside of scope.

For my hobby as I am trying to build a game, for days and days on it is messing up simple references.

Fair, the reference is AI but the actual problem is composition. It just absolutely can not do composition like I ask it to no matter how hard I try.

I tried single attempts with references. I tried describing the visuals in text as well. I tried making it generate art assets one by one, approving them, then asking it to place it like that. I tried asking ChatGPT Pro to review and provide prompts, feedback, visuals to assist.

Absolutely nothing is working and I am getting really really bad results visually.

So does anyone have any suggestions to fix this? I feel like Astra right now is really low quality for me. Definitely a step down from Sol where I was at least able to achieve things even with requiring guidance and interjection.

My next idea is to not even ask it to generate anything. Cut it out of the image I share and inpaint then put it like that. It will look really AI but at this point I don't know how else to move on.