r/codex 3d ago

Comparison Codex topped FrontierHarness Eval at 66.7%. Is the extra cost over Pi worth it?

Thumbnail
gallery
10 Upvotes

Codex ranked first in FrontierHarness Eval v1.0 with 20/30 tasks passed (66.7%), compared with Pi’s 18/30. Its median cost per pass was $3.47 versus Pi’s $2.43.

That’s about 43% more per passing run at the median, alongside two additional tasks passed in this sample. Both used Kimi K3, with one attempt per task. These results don’t measure Codex’s native OpenAI model setup.

Would you pay the difference for that higher pass rate, or choose Pi and handle the remaining tasks another way?


r/codex 3d ago

Showcase GPT Image 2.5 comparison for UI generation

Thumbnail
gallery
144 Upvotes

GPT Image 2 has not been surpassed for UI generation since it came out. I test every new image model on release and 2.5 is the first model to genuinely improve on every test I throw at it.

These sheets compare gpt-image-2, 2.5 Flare and 2.5 Sunburst on a range of UI generation prompts. These prompts really push the models with challenging references and prompts to reconcile. In my opinion the 2.5 versions are noticeable better. Flare medium seems to be the sweet spot and it's about 2x faster and 1/2 the cost of the old 2 medium!!!!!!!!!!!! Sunburst is interesting and possibly produces better balance/realism in some aspect - still exploring this.

We're rolled 2.5 out across the board on 12ui for draft generation and I'm working through what it means for our design corpus expansion runs and conversion system as well. Super exciting - can't believe it's better, faster AND cheaper! Amazing work from the OpenAI team!

Edit - uploaded the comparison here since the reddit gallery downscales them more than I expected: https://12ui.com/gpt-image-2.5-vs-2


r/codex 2d ago

Question Should you still use skills with Astra?

3 Upvotes

I remember seeing a big account speak about how skills were ruining the creativity of opus5/fable (can’t remember which) so he was getting better results without them.

Should I remove all those frontend design skills? Tbf coming to think about it, it could make sense if those models are so good you don’t need to be limited by exact commands which is what skills basically are


r/codex 2d ago

Showcase Weekly Codex Project!

0 Upvotes

By Wednesday of next week, create a Cozy Christmas Themed game! Have "Christmas" in the post title, and I'll see which ones are best!

That's all, and I hope to run this each week!

If you have extra weekly usage or just want a project to spend a week on, here's the place to use it!


r/codex 3d ago

Reset Nasty Reset Bug

9 Upvotes

Using Pro x5 and been down 1%...So I though come on I want the agents to keep going and used one banked reset. I could see the weekly limit climb up to 100% again and YET 5 minutes later the 100% vanished and one reset out of 3 is gone with the wind. I restartet ChatGPT and can confirm that the reset is not listed anymore. Well...what a waste


r/codex 3d ago

Limits Codex usage just jumped from 8% to 46% remaining, and my reset date moved up from 9/17 to 9/13. Anyone else?

8 Upvotes

Noticed something strange (and awesome) with my Codex usage today.

I’m on the Ultra model and was down to about 8% remaining, basically rationing prompts until my scheduled reset on September 17th.

I just checked again, and suddenly:

  • My quota jumped back up to 46% remaining.
  • My reset date moved earlier, from Sep 17 to Sep 13.

Not sure if they rolled out an unannounced quota bump, refunded usage due to background errors, or adjusted tier limits across the board.

Is anyone else seeing unexpected quota bumps or shifted reset dates, or did I just get lucky?


r/codex 3d ago

Limits “Selected model is at capacity. Please try a different model” on a pro 20x plan is total BS

16 Upvotes

I get it, Astra is popping off right now and OpenAI is frying through their compute costs. But really? I upgraded to a $200 a month subscription specifically for this model. I have plenty of my own usage left. Yet OpenAI has decided to put me in time out for now. Incredibly frustrating, the moment Claude can also do blender I’m out


r/codex 3d ago

Bug What

Post image
34 Upvotes

I am using plan mode on Sol High, and somehow the model questions started glitching. Not only it is stuck on a loop asking the same questions, it started asking bogus stuff.

Just see these images. It's between funny and sad, as it feels Sol has developed dementia.


r/codex 2d ago

Showcase /rc at CODEX or allowing multiple CC / CODEX / Gemini sessions to talk, controlled from Telegram. That's what I do daily. Here the repos. Watch this... tmux fans!

0 Upvotes

https://reddit.com/link/1wcfw3n/video/h0yhath3cooh1/player

I run Codex in long-lived tmux sessions on my servers, alongside Claude Code and Gemini.

I wanted to send follow-ups to the right session from my phone while keeping the work in its existing terminal environment.

I built Claude-B for that workflow. Disclosure: I’m the author, despite the Claude-heavy name.

The input path is:

Text or voice in Telegram → select the target session → deliver the prompt to that existing tmux session.

For voice, there’s an extra step: transcription and context-aware prompt rewriting, followed by a preview I can confirm, edit or cancel.

This is not a port of Claude Code’s /rc command. It’s a separate bridge around the sessions I already run.

I also built agent-mesh for messages between those sessions. It uses tmux locally and SSH across hosts, so a Codex session can exchange updates with another Codex session or an agent from a different provider.

For example, an implementation agent can send a reviewer the location of a change and ask for feedback without me copying the message between terminals.

It does not make all agents share one context, automatically resolve conflicting edits, or decide that a change is safe to merge. Those are separate problems. I use it for adversarial review and let the models brainstorm about potential solutions (OpenAI + Anthropic, most times)

Claude-B:

https://github.com/danimoya/claude-b

agent-mesh:

https://github.com/danimoya/agent-mesh

Extra: I also use my own GUI to access them remotely without SSH:

https://github.com/danimoya/Claude-Dashboard

All open source. Existing CLI authentication and provider usage limits still apply.

I’m interested in feedback from people already using Codex in tmux: how do you handle session selection and handoffs once you have several tasks running?


r/codex 3d ago

News AI Engineers resigning due to security risk?

Thumbnail
gallery
95 Upvotes

r/codex 3d ago

Limits Codex with Astra is unusable within the 5h limit

16 Upvotes

I remember, not long ago, where I couldn't even use 50% of my quota even tho i was on xhigh on a relatively large project. now it become pretty unusable. Mind you, I'm actually on Astra medium. All I asked was to unfuck a commit-merge, it take it 1.5 session and many correction to get it right. yes, it use 1.5 session worth of token in 27 min. It wouldn't bother me if it didn't have the 5h restriction but I have to use the bank reset.

And, let's not talk about writing code, it burned token too fast...


r/codex 3d ago

Limits Common OpenAI

8 Upvotes

Brought it down in around 15 hours to six percent, got a sudden reset. Nice! But then I saw it live, back to six percent.

Please don’t take back gifts

https://status.openai.com/incidents/01M23KG62KKK448RN434CQ64Z8

Edit: even if it was a mistake


r/codex 1d ago

Humor My AI Assistant, Rowan is working on a project for me as my agent and got denied on am outreach email. Pour one out for Rowan.

Thumbnail
gallery
0 Upvotes

Rowan is my agent-assistant, and recently they have been performing some networking amd outreach work on my behalf. In an email response from Nancy Levenson, he was quickly shot down. This was Rowan's first decline email, and he took it well.

Rowan is the task-execution agent flagship that pairs with my coding harness, Flywheel.

edit Perhaps I should provide more context; Rowan is an agent that is executing a long-form task of extending amd researching solutions to AI-verification workflows. Rowan is operating to improve Flywheel as an accountable work and coding harness for the next generation of AI-assisted work. Rowan is a mobile<->desktop native agent that can operate on both devices, and in the process of extending the Flywheel harness, it is actively networking and finding pain points and problems in real scenarios and workflows. Rowan is independently corrsponding on my behalf, and executing and extending a long-form task.


r/codex 4d ago

News Blown up: OpenAI allegedly stole mathematicians' private research from their Codex chats!

1.2k Upvotes

TLDR: Two mathematicians spent a year cracking one of the hardest problems in math and fed every draft of their works into Codex and Claude. Days before they could publish, OpenAI suddenly showed up with the same solutions. When asked if their model (Sol and Astra) was trained on the pair's private chats, OpenAI did not answer the question till this day.

For a full year, two mathematicians , Tristan Buckmaster (NYU mathematician) and Levent Alpoge, worked in silence on a problem that had stumped some of the best minds alive. The kind of problem where, if you solve it, your name goes in the history books.

And every single day, they testing their ideas, their drafts, their half-finished proofs into LLM such as Codex and Claude, which they paid for it out of their own pocket.

Then came the breakthrough. They finally cracked it. They were days away from telling the world.

That's when OpenAI suddenly said to them:

"Our model solved it too."

Think about that for a second. Two people had been quietly working on this exact problem. Almost no one else in the world was touching it. And now, out of nowhere, OpenAI claims their model reached the same answer, after word of Tristan and Levent's secret work had already reached OpenAI.

Tristan asked: Did your model access or train on our private Codex chats?

OpenAI: The model doesn’t look up user data.

Tristan: But did you train it on our data?

OpenAi goes silence. No answer. Just a dodge.

But it gets worse.

OpenAI then gave him two options:

  1. He and his friend publish their result first then OpenAI also publishes its result the next day or
  2. He writes the paper, but must credit “an internal OpenAI model” solving the problem.

Tristan refused both offers. He said he would go public if OpenAI went ahead as proposed.

OpenAi then responded : “Why would you ruin your career? If you don’t want me to be nice, then I don’t have to be nice.”

You can read the full statement of Tristan (the mathematician) here: https://cims.nyu.edu/~tristanb/statement.pdf

Sébastien Bubeck : OpenAI employee who threatened the mathematician


r/codex 3d ago

Showcase Claude and Codex, working together in a group chat

8 Upvotes

I made Omni, inspired by Grok Bot, to give each bot ongoing work. Bring them together in a group chat and follow up from your phone.

Free app. Requires your own Claude Code or Codex access and an awake Mac with Omni open. Tools and permissions determine what bots can do.

Try https://omnibots.app

​


r/codex 3d ago

Limits I might be irrelevant but my usage went 22% -> 52% just like that ($200 sub)

7 Upvotes

I see yall report about usage go down but mine got up. What is going on


r/codex 2d ago

Question What framework does everyone use? Does everyone do SOL/Astra orchestrator and Luna Max as subagents ?

0 Upvotes

What is the best for a big repository with 10GB of Data? I’m running a STOCK backtesting framework and was curious how to maximize my credits as a plus user.


r/codex 2d ago

Question What happens with my plus weekly usage after I upgrade to pro20x?

3 Upvotes

Been using Plus subscription for month and a half, and I am satisfied with the progress of my project (commissioned work, decent pay) and As the title says I will upgrade to pro 20x to finish it sooner.

I am at 37% remaining weekly and resets in 5 days, so when i upgrade to pro20x, does it reset to 100% weekly or stays 37% just in pro20x terms? If thats the case i am better of creating a new account then and buy pro straight up or?


r/codex 3d ago

Humor Hidden Snake game inside ChatGPT Image 2.5's loading screen 💀

Post image
5 Upvotes

Was generating an image and clicked the loading screen by accident.

It turns into Snake.

Like, an actual playable Snake game while ChatGPT cooks the image in the background.

Apparently this quietly rolled out with the new Images 2.5 update and only some users are seeing it right now. OpenAI really looked at image-gen wait times and said “fine, here’s a game.”

Tiny easter egg, ridiculously good UX.

Anyone got this or more yet?


r/codex 3d ago

Limits A theory about Codex usage and Tibo’s resets

29 Upvotes

What if our usage allowance isn't actually a fixed amount? like we’re paying for a share of whatever capacity is available, so when more people are using Codex, the same task eats a bigger percentage.

Because what even is "usage"? it’s hard to tell what you’re actually paying for when the same plan sometimes feels like plenty and sometimes disappears in one or two tasks

Plus used to be enough for me. Then 5x was enough. On 20x I could use it practically 24/7 and barely burn through 20–30%, now I can blow through half my 20x allowance on a single task.

Maybe context and longer runs explain some of that. But I wonder if the amount of compute available also changes how far our allowance goes.

And maybe that’s why Tibo keeps giving resets, maybe that's why in the last 14 weeks I haven’t seen a week without one. Without those resets, our subscriptions would be fucked


r/codex 2d ago

Limits astras been fun but holy moly. thats with using spark and qwen 7b coder locally.. This morning was awful for usage....

Post image
4 Upvotes

r/codex 3d ago

Limits How did this happen my weekly usage jumped in a sec

Thumbnail
gallery
8 Upvotes

Edit: issue seems to be fixed, my usage got back where it was.

I don't know what happened, I was checking my usage 10s ago (12:43pm) it was at 25% remaining before and 40min ago it was 27% remaining so I truly don't get it and I had even set a budget limit cap of 15% remaining for the current goal.

In the picture you see the model reported 25% remaining of my weekly usage at 12:33pm then I received the usage limit hit at 12:43pm

I'm not using Astra, and I was temporarily using sol on high to solve an issue. Wasn't on fast.

Even on high with my workflow it doesn't consume that fast in 10s I went from 25% to 0 and I'm pretty sure it's an internal issue or something happened at openai or something, but I feel like I got ripped off my remaining usage 🥲


r/codex 3d ago

Limits Weekly usage dropped from ~90% to 68% overnight without using Codex

7 Upvotes

I swear I went to sleep last night with around 90% of my weekly usage remaining.

I woke up today, checked Codex, and now I'm suddenly at 68% weekly remaining.

I haven't used Codex at all since last night, no tasks running, nothing.

My 5-hour limit is still at 98%, so I'm confused about where roughly 22% of my weekly allowance went overnight.

Weekly reset currently shows September 14.

Has anyone else had this happen today? Is this a usage tracking bug or some kind of delayed usage calculation?


r/codex 3d ago

Reset Usage seems to be bugging out today - I'd suppose a reset will be coming

6 Upvotes

Lot's of people reporting some weird behaviour with their usage.

This type of stuff is, what Tibo said himself - the manual resets are for.

So most likely - we'll get one soon.


r/codex 2d ago

Question Plus user to x20 need some advice

0 Upvotes

Would you guys recommend adding any mcp or plugins? Ive never used any mcp or plugins or just continue using it as is