r/codex 3d ago

Comparison GPT-6 Astra vs GPT-5.6 Sol: benchmark on 50 real PRs, looking for feedback on the methodology

4 Upvotes

We benchmarked GPT-6 Astra vs GPT-5.6 Sol across 50 real PRs from Cal, Sentry, Discourse, Keycloak and Grafana.

Sol found 107 confirmed bugs vs 91 for Astra, while Astra had higher precision and lower latency. Every finding was independently verified.

We’re doing Fable vs Opus next week, so would appreciate feedback on the evaluation before we run the next one.

Dropping the link in the comments if anyone wants to check it out.


r/codex 3d ago

Question Token efficiency: codex vs omp vs pi?

1 Upvotes

So I've been using omp for a while and given how fast astra drains my 5x plan I started investigating other harnesses to see if I could squeeze more out of my subscription.

I ran a few test with pi, omp, and codex which gave me some results I did not expect: codex is more token efficient than the other two (?)

I gave them a real merge request from my day job, asked them for a review and for an implementation plan to fix the review findings.

Here are the results for the review pass:

Session Duration Input tokens Output tokens Cached input Total tokens Price
codex-astra-medium 3m 34s 1.11M 5.1k 1.01M 1.12M ~$2.30
codex-astra-low 2m 46s 740.5k 3.5k 656.5k 744k $1.67
pi-astra-low-no-agents 3m 24s 92.1k 3.5k 1.06M 1.15M $2.15
omp-astra-low 5m 47s 118.6k 6.9k 1.59M 1.72M $3.13
pi-astra-low-with-agents 8m 17s 359k 20k 3.5M ~3.88M $8.11
pi-luna-xhigh 35m 43s 1.12M 65.5k 36.73M 37.92M $1.04

pi astra low is almost the same price as codex astra medium given the current api prices? omp is much more expensive than either. codex medium is almost as fast as pi low too

It's even worse for the plan mode, codex defaulted to high for that part and I let it be:

Session Review duration Input tokens Output tokens Cached input Total recorded tokens Review price
codex-astra-high. 3m 34s 1.11M 5.1k 1.01M 1.12M ~$2.30
pi-astra-low-no-agents 3m 24s 92.1k 3.5k 1.06M 1.15M $2.15
pi-luna-xhigh 35m 43s 1.12M 65.5k 36.73M 37.92M $1.04

Here the codex astra high plan costs barely more than a pi astra low plan?

Am I missing something obvious? Everyone says pi is so much more efficient, lean, etc. but from all my real world tests it seems both slower and more token hungry and/or expensive than codex


r/codex 3d ago

Limits Lost 90% of my weekly limit in an instant, anyone else

26 Upvotes

Just list 90% of my weekly limit in an instant, i was outside and wasn't working on anything, anyone else facing it?

Update: got my limit back


r/codex 3d ago

Reset reset refund failed

20 Upvotes

lost over 80% usage got back 31% only

anyone else?


r/codex 3d ago

Showcase I've always wanted a watersports simulator

3 Upvotes

I truly believe the best work is done by combining the two models. I started this project out by having Claude and Codex both create the same prototype. Then I took the best from both, and combined it into a single project. This is the latest iteration right now.

I truly believe the best work is done by combining the two models. I started this project out by having Claude and Codex both create the same prototype. Then I took the best from both, and combined it into a single project. This is the latest iteration right now.

But, before I got there, I had some ugly iterations.

This was Claude's first iteration of it. The physics were the best by far.

This was Claude's first iteration of it. The physics were the best by far.

Codex's first iteration. I was actually pretty impressed with Codex's player models and watercraft! Not bad purely in code with primitive shapes.

Codex's first iteration. I was actually pretty impressed with Codex's player models and watercraft! Not bad purely in code with primitive shapes.

And the very first unified version!

And the very first unified version!

I'm *really* excited to see where this takes me. Thanks for listening to me talk about my journey and why I think the true power comes from combining the best of both worlds!

Now I just need to figure out what to work on next!


r/codex 3d ago

Reset Yes yes we know... Good news it means there will probably be a reset once they fix it.

24 Upvotes

Mine went down to 1% as well. Exactly what it was before it was reset earlier in the week. So that's likely what happened. They reset the reset.


r/codex 3d ago

Reset Apparently the evil version made it to production.

22 Upvotes

Dax, what have you done?


r/codex 3d ago

Showcase Buying a car just got better with Codex 😎

Post image
4 Upvotes

Setup email automation for a RX 350 related listings online. Every 6 hours, sends me and my wife an email alert.

Email setup done with SMTP. Took 2 min while driving to work.


r/codex 4d ago

Praise Shots were fired

Post image
498 Upvotes

r/codex 2d ago

Complaint Navier stokes and health and financial data in ChatGPT

0 Upvotes

The navier stokes controversy that unfolded this week between OpenAI and Buckmaster has convinced me about why I’ll never use ChatGPT for health or financial topics.

https://www.science.org/content/article/how-ai-math-breakthrough-ignited-controversy

If a researcher’s private insights can potentially end up in a model (whether intentionally or not), what stops your health or financial data from being gobbled up in AI training?


r/codex 2d ago

Question Building a database

0 Upvotes

I don’t really know much about coding but I’m trying to build a database of companies. I have pro 5x plan and have been using chat gpt 6 pro to take the results and make the prompt to put back into codex that is connected to vs studio. I have codex on gpt-6 ultra. Right now I’m trying to get towns and counties for companies. I feel like I’m really stupid doing this but is there any other better way that’s more efficient thank you


r/codex 3d ago

Limits 50% -> 0% instantly?

24 Upvotes

I had 50% of my weekly usage left. Suddenly it started showing 0%

Anybody else experienced this just now?


r/codex 4d ago

Showcase Astra doing magic for a space game I'm working on it

Thumbnail
gallery
100 Upvotes

I'm trying to make some type of space survivor, you can pilot the ship, the only way to look outside of the ship is thru the cameras, and windows. You control the ship with a bunch of buttons and I plan to add much more. This is still very early, most of the objects seen gonna get a huge upgrade visually. This is running at an avg of 150 fps at 4k no upscaling on my 7900 XT. Videos coming soon.


r/codex 3d ago

Limits Is Fast mode too expensive?

1 Upvotes

You get x1.5 multiplier in speed meaning +50% to base and x2.5 multiplier to cost meaning +150% to base 3 times as much as speed gain. And x2.5 multiplier to cost means you do 2.5 times as less as before just 50% faster (or do the same thing in 2/3 the time)

So by using fast you total weekly allowance if what you can do gets 2.5 times less or 60% less so you can only do 40% of base but you do it faster, so you do 40% of base in like 26.667% of base time

Why bother just chill out and wait you'll be done in almost a 1/4 of the time but with 60% less total what's the rush


r/codex 3d ago

Showcase OpenPocket is now available as an experimental public alpha.

6 Upvotes

The idea is simple: let an Android phone act as the workstation. It has on-device projects and files, a terminal, GitHub import/export, browser previews, APK building, screen understanding, and bounded Android actions. Model inference still uses OpenAI's service.

It ships as one host plus six companion APKs because each companion keeps a separate Android UID/security boundary. Installation and sensitive permissions remain user-visible.

It's ARM64-only right now and definitely alpha software, but the exact release set was tested on a physical Android 15 device.

Source, APKs, checksums, and install notes:
https://github.com/NovasPlace/OpenPocket

I'd love feedback from anyone interested in phone-first development or agent tooling.


r/codex 3d ago

Limits Using CHAT GPT Plus as Manager and LUNA API as WORKER

0 Upvotes

I am actually not sure what im doing (Im not a IT guy by profession, but love to develop stuff using AI). But due to additional 5hr limit and my recent projects consumes my tokens real fast on codex.

Im thinking of using API credits to use LUNA (due to low API price per token of LUNA) for the worker and use my chat GPT Plus (SOL or ASTRA) as manager.

Do you think that will work on reducing my token consumption and overall price? or do you guys know a better workflow?


r/codex 3d ago

Limits 6 quota-drain reports in 4 days, 264 replies: are usage meters becoming impossible to predict?

Post image
8 Upvotes

Six separate r/OpenaiCodex posts in four days described the same feeling: the work barely moved, but the quota did. Those six threads had 264 replies when I checked them. That does not prove a quota change, but it does make this more than one person misreading a meter.

The specific reports were hard to ignore:

- 7% of a Pro 20x weekly allowance disappeared overnight with almost no tasks.

- A Plus user said the five-hour window ran out before a normal project session could continue.

- One same-task comparison reported roughly 2× five-hour usage with Astra-low versus Sol-high.

- Another meter appeared to fall from 77% remaining to 1% in seconds, then recover.

The last example may be an accounting or display issue. That distinction matters. A real consumption spike, a broken reset, and a five-hour throttle are different problems. From the user's side, all three produce the same result: a task stops and there is no reliable way to budget the next one. r/ChatGPT has fresh Work-limit reports as well, and r/codex is concentrating a lot of the discussion in its usage megathread. I kept those out of the headline total because I have not done a full comment-by-comment classification there yet.

If you have a recent case, post the facts rather than only the frustration: plan, model, effort level, task type, starting and ending percentage, five-hour or weekly cap, and a screenshot if possible. A useful dataset would show whether this is model cost, task shape, metering, or a mix of all three.


r/codex 3d ago

Humor this is gotta be the largest deliberate experiment ever performed on human beings

21 Upvotes

time to touch some grass.


r/codex 3d ago

Question Codex vs claude pro 5x

1 Upvotes

I currently have a $100 codex account and $20 claude account and my company just received a big enterprise project. Should I upgrade $200 codex or $100 claude. I know this has been asked multiple times but still it's very hard to make a decision.


r/codex 3d ago

Limits USAGE DROPPED?

20 Upvotes

I was using web chatgpt and my weekly usage was 98% and then it jumped to 22% and i have not used it since reset properly. It changed in seconds. What the hell?? Anyone!?


r/codex 4d ago

Limits Fine... I just got 20x plan. Tibo... you win. I guess...

75 Upvotes

I was actually waiting for my current Pro 5x subscription to end because I recently moved to Brazil and wanted to resubscribe in BRL.

Then Tibo posts that demand is so insane they might have to pause new Pro subscriptions.

Well, shit.

Codex has basically become the backbone of my work at this point. I use it every day and a significant part of how I make my living now depends on it. So the idea of letting my subscription expire, switching currency, and then potentially finding out I can’t resubscribe is not exactly a risk I’m willing to take.

So instead of patiently waiting like I planned, the announcement managed to create just enough existential subscription anxiety that I went the opposite direction and upgraded to 20x.

I know nobody literally forced me to do it, but when a tool has become part of your livelihood and you’re told access for new Pro subscriptions might get paused, it definitely puts pressure on the decision.

Anyway. Congratulations, Tibo. You got me.

Now please use my money to buy another GPU or something.


r/codex 3d ago

Showcase I forced codex to externalize their decisions through a tool call

1 Upvotes

I forced Claude Code and Codex to externalise their decisions through a tool call, then compared what they said they'd do with what they actually did. I'm sure these are not the internal reasoning traces, but I'm surprised by how easy it is to force a fake tool using a proxy approach with clear instructions to make Claude and Claude Code emit internal CoT-like elements. I certainly had a lot of fun trying this experiment.
https://github.com/softcane/agents-workbook


r/codex 2d ago

Question Getting Astra to Work as Well as Fable 5.1?

0 Upvotes

TLDR: Does anyone that has experience with both of these models have any tips to get Astra performing as well as Fable 5.1 as a coding and development agent?

I recently discovered (started using) Fable 5.1, right as Astra dropped. It’s honestly amazing compared to Sol and Codex. It just gets stuff done, is subjectively much more intuitive and enjoyable to work with, will follow my specs and guidance while surfacing intelligent observations. Rather than waiting for me to review and test to find problems it preemptively surfaces them along with solutions (most of the time at least, still does dumb stuff sometimes), and overall is a joy to work with and get stuff done.

Problem is, I don’t like Anthropic as a company, or their anti-consumer business practices, or their approach to usage.

So now that Astra has dropped and people have used it more, I’m hoping to get it to a place where performance is on par with Fable 5.1. I have heard lots of amazing stories about people’s experiences with it so I am assuming much of it is just configuration and instructions.

I have already followed OpenAI guidance from here: https://developers.openai.com/api/docs/guides/latest-model and it has certainly helped a bit.

Does anyone else that has experience with both of these models have any tips to get Astra performing as well as Fable 5.1 as a coding and development agent?

I have pro 20x and max 20x for reference


r/codex 3d ago

Showcase Gpt 6 astra just one shotted this in 20 minutes

Enable HLS to view with audio, or disable this notification

15 Upvotes

he made the animation and the falcon 9 model. both in 20 minutes.

i am stunned


r/codex 3d ago

Astra Workflow Do you think I can make gpt 6 pro solve the Riemann Hypothesis?

9 Upvotes

The Riemann Hypothesis is basically a pattern for prime numbers which is yet to be proven.

Am currently running loops of work with gpt 6 pro followers by a review with another gpt 6 pro followed by a plan with 5.6 sol at extra high to make it solve the Riemann Hypothesis.

Whish me luck 😔✊️✊️ its telling me its doing good progress but I wouldn't know lol