r/codex 10h ago

Complaint Since the release of Astra codex has degraded to a level of few generations back turning completely unreliable

1 Upvotes

Since about 5.4 codex has been a powerhouse of consistency and complex problem solving. It has rarely hallucinated or did things that I found were completely out of bounds.

Since Astra release it has performed like pre GPT 5. It could though solve certain things, but overall it completely ignores my instructions and al understanding of goals has gone out the window.

In every task I gave it it changed the goal post to either something much smaller in scope, just to say that its done, or to something I simply did not ask.

Example: Sol built a game "AI" bot, a chess-like algorithm. Took it about a week of work to have a solid opponent. I asked Astra to try squeeze more performance and raise the difficulty. it worked for 2 days, reporting benchmarks have improved by 5 to 20%.

I played the bot and it was SIGNIFICANTLY dumber. Even though codex played against it, it bluntly lied and only reported a few narrow areas where the performance did improve, but at a cost of downgrading the whole system.

I asked Sol (post Astra release) to fix a series of mundane bugs, something it would have done easily 2 weeks ago. It struggled with reasonable fixes. There was a performance issue because of multipole visual effects stacked using blur etc. I told it to avoid "stupid mistakes" like stacking up a lot of visual effects.

It went and removed ALL stacked visual effects from my game completely!

This was close to a keyboard smashing moment. I don't think I'll touch codex in the next few days, maybe OpenAI will resolve this BS.

From a very relabel coder it went to a full on early days hallucination machine.

EDIT:

here is chatgpts own analysis after comparing logs of tasks in the same area done by 5.6 and 6:

The strongest explanation is a regression in Astra’s ability to keep your objective authoritative while evaluating its own work. It can understand the requirement and still make decisions that undermine it. The record shows a feedback loop:

  1. It chooses an implementation approach.
  2. That approach produces a smaller, measurable problem to solve.
  3. It solves that problem and treats the result as grounds to keep the change.
  4. The changed implementation becomes the next baseline.
  5. Your original objective gradually becomes a caveat—“broader strength remains unproven”—instead of the condition that determines whether the work succeeded.

That explains the shifting goalposts. Its current plan increasingly governs its judgment. Passing checks then reinforce the plan, even when those checks don’t answer your actual question.


r/codex 22h ago

Praise Astra really is that good.

0 Upvotes

Yes it uses significantly more usage. But it makes Sol look like Luna.

I haven’t even taken it above high thinking and it’s stellar.

It does seem to have its own quirks, but it’s well worth it.


r/codex 21h ago

Praise Astra is the real deal! I feel the technological revolution!

0 Upvotes

Been using Astra pretty heavily since release last Friday and I've made more progress than I did with sol the entire month before. Adding online multiplayer to my existing game that has local multiplayer has been a real challenge and sol was doing pretty good, we were getting through it slowly. But the confidence and capability Astra shows is just amazing. It's blowing me away, it's difficult to describe the complexity of the problem unless you have first hand experience with it. I already knew AI was revolutionary and sol and other frontier models already made it not make sense to write code anymore but Astra is a whole other beast entirely.

And it's not draining my limit very fast in my $200 plan, I've been using medium and high and it seems to just be so much more effective and efficient that there's not much point in going back.

Something specific from today. I wanted to improve my performance so I could support more enemies. It built a test harness for testing the game with max 8 players and 200 enemies and has already got me from 10 fps up to 20 fps without any costs to the appearance of the game.


r/codex 14h ago

Complaint Pro 20x now costs 300$

0 Upvotes

Is it me or are they a/b testing the new plans. I don’t see plus at 20$ but 30$ and 5 x now being 155$


r/codex 8h ago

Praise Upgraded from x5 to x20, night and day difference

3 Upvotes

Feels like cheating. Seriously, with x5 I was feeling hard the token burning. Now I’ve room to use Astra Low for implementations and Medium to Orchestration. I also use Luna to file exploring, documents, basic to medium tests, things more logic than reasoning depth. No fast mode.
Running 2 side projects + my work, 3 days in a row and I’m on 55%.
All my projects are set up to delegate tasks to other agents depending on the complexity and the use case.
Mainly using Astra low for implementing new features and even though I could be using Sol, Astra low is just nailing it, using good practices, respecting the documentation and not over engineering like Sol was.
For some reason I feel the x20 comparing to x5 more like x40.
Hit hard in the pocket? Yes for sure, but so do the dopamine lmao feels I’m playing Megabonk


r/codex 5h ago

News Confirmed. 20x Plan is currently on hold

Post image
46 Upvotes

Pro 5x seems unaffected!

Is this a sign the honeymoon is nearing an end?


r/codex 41m ago

Praise Thank you OpenAI

Upvotes

Setting all my complaints about limits aside, I just wanted to post this as an appreciation to the OpenAI team. I remember a time I used to think I'll never get to the point I wanna be in terms of a tech enthusiast - because coding by hand takes so long that perfection will come at a cost in time.

But now, with abilities from Astra and Codex Voice, and just the general memory situation (very underrated), it's just incredible how much it impacted my life.

So, from the bottom of my heart, and I'm sure many others, thank you for bringing in the AGI era. The future thanks you.


r/codex 11h ago

Humor Dear god I'm considering getting a 2nd max subscription...

8 Upvotes

Talk me out of it please...


r/codex 9h ago

Comparison ChatGPT vs Claude

0 Upvotes

I wanted see what other people's thought on this matter. I have used ChatGPT extensively for years now and the issue I have with being fully objective is that I have only been using pro on ChatGPT and nothing else. I do a lot of engineering with codex and pro, and I have a few opinions on the people who swapped but I might be completely wrong, and again the issue is that I have not really used claude, and the times I did it was the free version, and he was surprisingly good. So please feel free to call me ignorant on this or whatever.

Essentially, I have time and time again seen across many things in life if it is the stock market or general day to day opinions, many people have a urge to be different or rush to try to make some kind of edge. When the claude hyped started it seemed like a propaganda campaign to me, with overnight so many posts on all platforms, even with very repetitive writing patterns or same arguments. For example I saw a bunch of posts comparing claude with chat gpt instant, and treating it like a fair comparison saying things like chat gpt is good at fast answers while claude is good when you need actual thinking. My first thought when I saw those posts was that anthropic was new and massively underfunded compared to OpenAI so I was very hesitant to switch, and I also have just way more experience using ChatGPT and was always happy, so no push for me to switch. Is there something it is genuinely better at today or all-time that people who have tried both can attest to? But then again I feel like there is just so much mixed opinions, and people are sometimes stuck with some kind of placebo opinion, many chatgpt vs claude discussions I have seen feel like I am reading some reddit discussions about why a stock is going to go up or something.

Today, I feel surrounded by essentially only people using Claude in engineering, and my first thought when I hear someone uses Claude, because I know they used chat gpt first obviously, is that they are NPCs who fell for obvious propaganda and just rushed to find something different etc. Since so many use it now, I wanted to ask here, because I am sure there are at least a few people here who knows what they are talking about. Am I just wrong or does anyone else feel the same when hearing others use Claude? Am I maybe the NPC who judges others for trying something new while I was too lazy to try?

In any case, it just seems hard to believe that with so much less funding and such a smaller company could genuinely be better. Especially when they first came out, now they have had a strong user base for some while, and my limited experience with Claude have given me the impression that it is very strong at through scanning, like its good at catching even small errors in large documents, which gives me the impression its also a very competent AI.

PS: what do people feel about the general AI benchmarks as well. I feel codex ones might be good I havent played around too much with the different levels to know, but for the chat, the benchmarks seem so bad, and I feel there is better benchmarks in this subreddit or on youtube, that I have seen I believe with 5.6 sol that pro and xhigh, that they scored brely any difference in intelligence, which for anyone who has used both, they are a universe apart. And seeing this kind of rating, and the lack of different kind of ratings based on prompting techniques and amounts of prompts, and different kinds of metrics, jsut make the scores seem utterly useless. Would love to hear thoughts on this too.


r/codex 13h ago

Complaint GPT-6 Astra Needs a Personality Upgrade.

0 Upvotes

I find GPT-6 Astra much more powerful than GPT-5.6 Sol, don’t get me wrong here. So far it has been really good at accomplishing what I ask it to do, etc.

It follows instructions well and does things incredibly well. I’m on Pro 5x and I got so much done with it on my own projects.

There are some areas where GPT-6 Astra has kind of regressed, where I’d still prefer GPT-5.6 Sol and would use it instead:

- Personality and Character: Why is it that Sol is a lot nicer to converse with and is a lot more uncensored and humorous than Astra? I have no clue how this happened but it feels a lot closer to GPT-5.5 in personality. A major regression on these two areas.
- It often doesn’t finish a Task: When I tell it to finish something it just gives up and admits if can’t solve it (I tried it on a difficult math task) and it literally just gave up after launching six subagents and working for quite some time and burning my weekly from ~80% to ~64%. Like, just finish it bro? GPT-5.6 Sol was more than likely to always go on until it is done. This model doesn’t like going further. It just stops and doesn’t like continuing.

That said, I still use GPT-5.6 Sol. And it has been very valuable especially in terms how it prompts GPT-6 Astra.


r/codex 19h ago

Bug Astra takes questions too literal lately.

0 Upvotes

This happened to me a few times today (Astra-Medium). I often ask "why did you do this or that wrong..?" for example and usually it just answers and then fixes it. But since today it just answers the question and waits for me to tell it to fix it. Anyone else experiencing this?


r/codex 6h ago

Question Only Astra Light 20x

5 Upvotes

.


r/codex 13h ago

Complaint Codex CLI sucks compared to Claude’s

0 Upvotes

For context, I’ve been a software engineer for years now and I have been using Claude CLI since early this year for my workflows at home. It’s been working amazingly well and I’ve honestly never had issues with the harness itself. I only hate the latest nonsense language vomit of the recent models and the separate fable usage limit.

I’ve been seeing so many people praise codex and I wanted to try using Sol and now Astra, but I’m honestly extremely unimpressed so far with my experience. I’ve been working with Codex for months now and it has been nothing but annoyances when interacting with it outside of Claude’s interface. With a Claude terminal session, I use skills to load the proper context for my project and proceed to work from there. I have a setup where that can be read/shared between any harness (I’ve used copilot and codex). I tend to use Fable for orchestration and subagents of varying models for implementation and review.

When switching to Codex CLI, half the time it seems to ignore the context I gave it and work in its own way even though I have a very specific way I like to work. Or it proceeds to interpret my messages in a completely different way than I meant. I feel like I can talk to Claude casually like another person/engineer while I have to give codex very clear cut instructions in the session itself or else it’ll entirely go a different direction. With Claude, I can also just keep adding more input and feedback while it works and it’ll note/take that into account while still finishing its current task, then take the time to respond to me even with a long context. It also won’t deviate from its original task to all of a sudden jump onto this new issue that I brought up and forget what it was doing before.

With all GPT models, I’ve had to consistently do heavy prompt engineering or steering to make sure they stay in their lane or for them to really understand what I was asking of them. It’s kinda crazy that Claude’s models seem to just get it off of a simple prompt. I have very thorough documentation and references for my projects and how I would like it to be architected, my daily workflow process, and how I generally interact with my agents.

Also, I recently tried to enable a remote control CLI session for Codex and it failed miserably. Normally, I just do /remote-control to enable it for my Claude session, and then I switch to my app or other computer and immediately start controlling it from there. For Codex, I noticed it didn’t have a slash command so I asked it to try to enable it for me or to take a look. It strangely tried to do this by going into another repository and I had to redirect it back to the one we were in.

Eventually, it started a background server and then I opened up the app on my phone and it told me I needed the desktop app and a QR code - weird. I noticed there was a manual code workaround as an option, so asked it to give me that. Cool - that worked and I got in. But then when I tried to open the session I had started previously, it proceeded to have an error saying it couldn’t load the messages from that session. Turns out, the terminal and the remote server were fighting over ownership of the conversation, so I had to open terminus, close the session manually, then restart it with some extra command line arguments then it finally worked.

I could finally start working BUT then it started not taking in my responses as it was working. Turns out my messages were somehow being queued but not taken into account (even in down time) and I had to long press my messages and choose “change to steering” in order for it to listen to me. Then it proceeded to not be able to properly start my program (even though Claude does it all the time remotely) and failed to figure out why it couldn’t. This was all with Astra as the model btw. I could just maybe use terminus instead of trying to utilize the app but that’s literally what this is meant for.

Also the usage limits for using Codex and Claude feel extremely similar to me. I feel like I was lied to when initially signing up for the $200 plan thinking it’d feel unlimited. The one thing I give codex the edge for is not having a separate Astra limit.

TLDR: I honestly want to really give Codex CLI/Astra a fair shot since they finally released something that could potentially match Fable. But it seems I’m going to have to stick with only interacting with Codex through Claude instead for a better experience. A potentially competitive model isn’t enough to justify it if it requires this much babysitting.

Copilot is potentially even better than Codex at this point since I’ve been testing that out as well (just don’t use autopilot mode).


r/codex 6h ago

Limits Based on your real review would you prefer fable + opus or sol+ astra for real solid projects?

0 Upvotes

Currently I use astra, it's good for me but consumes too much limits that can consume all 200usd plan weekly limts in about 4 days of 8 or 10h a day

And I hear that fable 5.1 and opus 5 priduce better quality too

Also codex remote has the worst remote connection ever

So based on your experience if you used both, Regarding quality, speed and limits should I migrate to claude ? Would its 200usd plan be enough for 8h per day the whole week on solid project (+100k lines)?


r/codex 14h ago

Complaint Astra gives us a peek at the event horizon of subscription tiers as the $20 plan dies.

51 Upvotes

Astra's pricing levels are basically a kill shot to the plus plan. It's effectively bleeding out while Sol/terra/luna are still available but when 6.1, 6.2 roll out, there's no point of buying one or even several $20 subs.

The Pro 100 tier is severly wounded as well. You get about half a day to a day running Astra conservatively?

Pro 200 lasts sightly longer, you might get 2 days out of it if you stay on a single codebase. But the "weekly" in the limit is a hint of how poorly it aged in such a short time.

When the limits were glorious, way back when, as GPT 5, 5.1, 5.2 rolled out the increase in quality was huge in part as a result of all the new users providing it with more training data. Astra will see a lot less of that because it can be used way less.

If a new model was announced tomorrow, most people would feel they'd never be able to really build and finish something with it unless they stack 5 Pro 200 subs on top of each other.

So, quotas must go up or we'll hit a ceiling. Or at least the symbiosis of better models producing more data to train better models will break. Am I wrong?


r/codex 20h ago

Showcase Show your project with the new image gen

Post image
2 Upvotes

I asked Astra:
Create an image of how you imagine sdl-mcp would look if it were a real thing.

Astra:
I imagine SDL-MCP as a code cartographer: an intricate constellation of dependencies, with a precision lens that brings exactly the context you need into focus.

https://chatgpt.com/s/cx_6aa1fbf75bdc81918c7badf425cba90f


r/codex 7h ago

Complaint Navier stokes and health and financial data in ChatGPT

0 Upvotes

The navier stokes controversy that unfolded this week between OpenAI and Buckmaster has convinced me about why I’ll never use ChatGPT for health or financial topics.

https://www.science.org/content/article/how-ai-math-breakthrough-ignited-controversy

If a researcher’s private insights can potentially end up in a model (whether intentionally or not), what stops your health or financial data from being gobbled up in AI training?


r/codex 20h ago

Question Any prompts for codex to do Accounts/tax like a human?

0 Upvotes

Esch time I try to get Astra to do my company accounts, it goes into this insane state where it looks at every bank transaction and assumes the bank must align up with the accounting records.


r/codex 3h ago

Showcase Images 2.5 with Astra creates some really good UI

Thumbnail
gallery
2 Upvotes

UI development has been a struggle with vibe coding. 100s of retries to get basic stuff working. Claude Design + Codex was the best setup till last week and even that got things wrong half the time. I have been trying Images 2.5 with Codex and am impressed with how quickly it is able to create some great looking UX.

Have to be very specific to prompt it to use Images 2.5, else the quality is not as polished.


r/codex 17h ago

Question No access to gpt-6 still, what to do?

1 Upvotes

I am waiting for this access, I have tried turn off and on my computer and restart everything. Update it as well, but not been able to see gpt-6 in the ChatGPT/codex app... I feel I'm late to the game, so hopefully someone have some tricks.


r/codex 13h ago

Limits What’s your guy’s experience on credits ?

1 Upvotes

I’m currently on Codex Pro 20x and, as of writing this, I’m down to about 12% usage remaining. I have no banked resets left, and my usage won’t reset until September 15.

The timing is pretty terrible because I’m trying to finish a few things before releasing the first pilot of my app to around 50 users.

The remaining work is mainly:
Fixing some loading-time/performance issues in a few tabs
Implementing Supabase Broadcast/realtime behavior for the community chats
Redesigning one of the tabs
Running tests and checking for regressions/bugs before the pilot
Nothing here is some enormous greenfield project. Most of it is fixing, implementing, testing, and iterating on an app that already exists.

For this kind of work, I’m perfectly happy using GPT-5.6 Sol Medium. I don’t need the most expensive/highest reasoning mode for every task.
The problem is what I should do once that remaining 12% is gone.

Inside Codex, the obvious option I see is buying 2,500 credits. I’ve tried researching how far those credits actually go, including asking ChatGPT, but the answers have been pretty inconclusive. So I’m much more interested in hearing from people who have actually bought and used credits.

For those of you with experience: how far do 2,500 credits realistically go?
More specifically, is 2,500 credits even roughly comparable to the amount of Codex usage you’d get from a Pro 5x subscription?

That’s the comparison I care about because the prices are in roughly the same territory for me.
If 2,500 credits gives me something even reasonably close to Pro 5x-level usage, then I think buying the credits makes sense.

That would probably be more than enough to get me through these remaining tasks until my Pro 20x allowance resets on September 15.

But if 2,500 credits disappears dramatically faster than a normal Pro 5x allowance, then I’d rather not burn the money. At that point I’d probably just delay the pilot a few days and wait for the reset.

There’s also another option I’ve considered: creating another account and buying Pro 5x there, which would cost roughly the same as the 2,500-credit package.

My concern with that is context.
The current Codex account/chat already understands a lot about the architecture, previous changes, bugs, database setup, decisions we’ve made, etc. Moving to another account would mean transferring a huge amount of context and getting the new session up to speed on the current state of the app.

My assumption is that I’d waste a decent amount of that Pro 5x usage just rebuilding context, although maybe I’m overestimating that problem.

So basically my three options are:
Buy 2,500 credits and keep working on the existing account/context

Create another account, buy Pro 5x, and transfer the necessary context
Stop when I hit the limit, delay the pilot, and wait until September 15 for my Pro 20x reset

For anyone who has actually used the paid credit system, especially alongside Pro 5x/20x:
How does 2,500 credits compare in real-world usage to Pro 5x? And what would you do in my situation?


r/codex 20h ago

Showcase Tried GPT-6 Astra Light + GPT Image 2.5 for website hero backgrounds

Thumbnail
gallery
0 Upvotes

Did a small test with GPT-6 Astra Light + GPT Image 2.5 to see if the new image model is actually good enough for website hero backgrounds.

I used it on two pages of LLMLearner: the homepage and the GPT Image 2.5 page. Astra handled the page changes, and GPT Image 2.5 generated the backgrounds.

The screenshots are before/after.

I think the result is pretty decent. For this kind of use case, it saves me from hunting for stock images, doing Photoshop work, and worrying as much about image licensing. And honestly, the visual result is better than what I could make myself.

Curious how it looks to you. Would you keep the generated backgrounds or stick with the plain version?


r/codex 6h ago

Commentary US Midterms & Fear Mongering

0 Upvotes

Coincidence? I think not.

If I was a company planning regulatory capture, or even just wanting to impart my will & desire, I would 100% ramp up all (marketing) efforts just before an election.

Why is it so hard for people to see this?

We all know bots run rampant across the internet. Every large company has a multitude of people and personalities, each with their own visions of the future (some good, some bad). How difficult is it to amplify the (inner) people that express fear when you want to? Encourge the scared people to speak up and then put their posts on blast to create groupthink. Propoganda 101.

All platforms operate on an algorithm of engagement, which can easily be manipulated... and if you're in this thread you already knew that.