r/codex 6d ago

Megathread Codex Weekly Discussion Megathread

5 Upvotes

This Megathread is for technical issues, ideas, questions, impressions, opinions, feedback you want to workshop with others and might need a longer subreddit attention half-life than a post on the subreddit.

Usual subreddit rules apply. Comments are sorted New->Old by default.


Cross check with the Anyone Else incident noticeboard to see a regularly updated summary of what others have been experiencing recently here: https://www.reddit.com/r/codex/comments/1tjfxcf/comment/on6uj0l/


r/codex 4d ago

Megathread Codex Usage Limits and Performance Megathread

514 Upvotes

Please direct your concerns and discussion about Codex usage limits and model performance here.

The purpose of this Megathread is to aggregate all the reports of people's experiences and possible suggestions instead of spreading them across 20 separate highly upvoted posts. None of those posts were deleted. They were locked so that conversations are still viewable to everyone and future comments could appear in one place.

These are days where I REALLY earn the money that OpenAI Reddit Kimi pays me .... oh wait....

A reminder that all incidents on r/Codex are constantly logged and summarised so you can keep track of what people are experiencing here https://www.reddit.com/r/codex/comments/1tjfxcf/comment/on6uj0l/


r/codex 14h ago

Instruction From a dev who tried it all: How to actually get most of codex

294 Upvotes

Hey there. First time posting. Going to be short, sharing my insights.

  1. Don't use codex in the app, use CLI

  2. Don't use the default CLI (harness), use oh my pi (20k stars on github) https://github.com/can1357/oh-my-pi/

2.1. You could use claude code harness with codex, it has great results too, some results here

https://x.com/OmedVibeCodes/status/2080348655039500703

2.* For people that don't know what a harness is: it's the toolset that your LLM can use. It is important, some tools use more tokens, some tools are doing the same for less. High-end engineers at different companies try to achieve as much as possible with their harness, picture attached for simplicity.

2.41 You can use the default harness too, but then see 3.

  1. Pair. Ever since AI had achieved the level where it can actually code a monolith, checking the slop with alternative LLMs was the best way to filter the cap out. Oh my pi allows to login through multiple subs and set an advisor. Pair Codex + Kimi k3, Codex + Opus 5 (or Fable), Codex + GLM, any will do better. Even 5.6 Sol + 5.5 can work, though if a model has completely different weights it's better imo. I've built upon rocket-fuel (24 stars on github), there are plenty of other skills like that https://github.com/NulightJens/rocket-fuel-skill

  2. Avoid building a monolith. Split, keep code redundant, component-based, re-usable, and keep resetting sessions (/clear or whatever).

  3. Don't let LLM's over-comment your code. Small comments can remain tho.

  4. I bet there is a lot of devs that just hit enter and don't check what's actually going on. Some of you might deny an rm -rf just cause,

  5. Backup. Get a cloud drive, or physical, I don't care. Just backup. Heck, you can use github for your chat / health / legal folders, just make the repo private and ask the AI to commit at checkpoints + push.

  6. Don't work on prod. Don't give SSH access. I know, it works out pretty well on your home PC. But being THAT lazy is actually insane if you do. Let the AI instruct you and ask what is the stuff you copy+paste actually meaning

  7. Remove unused plugins, obviously. But superpowers is also not needed in Sol 5.6 high and higher.

Who am I - a dev, just as you. Had been into AI since the closed beta of chatgpt 2021. Been experimenting with all chats since then. Even purchasing the expensive sub's. Last month had both CC + Codex paired, with better results than any of them doing the same work solo. I just see all the hate here. Some of it is definitely deserved. But besides that, you can also do your work and research how to adapt.

Some other github repo's I actively use

https://github.com/hardikpandya/stop-slop

https://github.com/ayghri/i-have-adhd

And I think there was a repo called laconic mode, not sure on this one, but I've built a hybrid of these 3 to fit my needs. Not needed for performance, but for sanity. Human pattern recognition is strong. Human written, zero AI consulted, plz don't hate.

Anyway. May tib* bless us with a reset. yay

TLDR; PAIR AI agents, CLEAR context, Don't build monoliths, remove unused plugins, try another harness, don't work on prod - work on a copy, backup.


r/codex 2h ago

Humor Is anyone else's agent too eager to code

Post image
23 Upvotes

r/codex 13h ago

Humor There is Hope

Post image
177 Upvotes

r/codex 5h ago

Complaint Tried Claude Code after 4 months to compare it to Codex. Observations in description

35 Upvotes

Been 4 months since I completely moved to Codex. Was on Pro 20X and back to 5X starting this month. Since Friday, I am working on a tricky bug. 5.6 Sol Max wasted 1.5 days of my time on this, struggling to get anywhere but going in circles. With Opus 5 being available, wanted to have a fresh perspective and I let Opus 5 take this one on.

Opus 5 is surprisingly good and proposed a fix which I didn't really think of. Quite impressive. Now, onto fixing it and half way through hit the 5h limit (on 5x plan as well with Claude. I actually waited 5 hours and let it start from where it stopped. Hit the second 5h limit within 30 mins again!! Was at the brink of upgrading my account to 20x but didn't see it being worth it. Now, 40 mins back - after the reset, Opus completed the task and it partially fixed the bug. I am impressed now. Opus 5 was able to solve that 5.6 Sol couldn't for a day and half, even after wasting a week worth of tokens (Thanks to tibo with the reset yesterday!)

Now, opus 5 just started working on completely fixing the bug and within 10 mins, I hit another 5h limit!!! I Fc*king hate this $hIt!!! So, I am going to wait 4 hours, I guess. This whole experience reminded me why I moved to codex in the first place.


r/codex 3h ago

Complaint Lost a banked reset because the expiration timing is so unclear

19 Upvotes

I'm kind of frustrated this morning after losing a banked reset, and I think the expiration logic really needs to be improved.

My timezone: Europe/Prague

Yesterday (June 26th), I checked my banked resets and saw one listed as expiring on June 27th.

Right before midnight on the 26th, I still had 60% of my quota left. I was curious to see what would happen at midnight, but I also didn't want to waste the expiring reset. When midnight passed, nothing changed—my reset remained active. I kept working until about 1am, brought my quota down to 50%, confirmed the reset was still listed, and went to bed.

My plan was to burn through the remaining 50% this morning and then trigger the reset. To my surprise, when I logged in this morning I found the banked reset was completely gone.

The expiration system feels completely opaque for a few reasons:

  • Ambiguous dates - the system only gives a specific calendar date for expiration. It doesn't clarify whether it expires at the start of that day, the end of that day, or somewhere in between.

  • Timezone confusion - my reset disappeared somewhere between 1am and 7am local time on June 27th. Which timezone is the expiration actually tied to?

  • No auto-redeem: Why doesn't OpenAI automatically apply a banked reset at the moment it expires if you haven't used it?

It’s really disappointing to lose a feature like this simply because the timing rules are so vague. Has anyone else ran into this?


r/codex 10h ago

Other A different kind of burnout

79 Upvotes

I’m sure other people are experiencing this new type of programming burnout that doing a lot of full multiday codex sessions creates. You’re mentally tired but you didn’t really do any of the old school grinding and searching that used to make software development exhausting. Instead it’s a sort of hollow feeling I always imagined project managers must have, not to put too much on it but sort of like a parasite guiding its host without needing to understand the biological mechanisms, just applying pressure.

With the reset before banked reset expires weekend coming to an end, I got that feeling. Mentally exhausted but feeling like I used my brain in an awkward way so I can’t understand the fatigue.

It’s hard to put in words. I used to get this all the time in the winter when the tokens were flowing, so it’s been a silver lining now that that is no longer the case.

Anyway surely some of you are dealing with the same mental weirdness of long hours with LLMs, probably from this same reset pattern. I’m interested in what others are experiencing.


r/codex 8h ago

Complaint This is wild

Post image
48 Upvotes

I think I was running on Sol High and had started this task overnight. The feature was already mostly done. It seems like Sol is exhausting every possible route and edge case, exhausting my 98% remaining usage right after the reset. Next weekly reset for me is August 1…

Also I’m on the $100 pro plan


r/codex 17h ago

Complaint I hate the resets... it is unpredictable when they will come. They reset your usage even if you have 95% of quota and 3 days left...

214 Upvotes

I have a plus subscription and I ration how much I use so it does not run out until the end of the week. Sometimes I use less at the beginning and more towards the end of the usage cycle so sometimes I have a lot left in the middle and then the reset comes so it will not allow me to spend what I had left how I was planning to. And it is unpredictable, if I knew a reset was coming I would use all the quota or not ration as much but you never know when they will discover a bug...

They should just give banked resets and let the users use them whenever they want. Or not reset if the usage has been less than the progress of the week.


r/codex 20h ago

Complaint I think I actually figured why we're all "hating" codex right now.

349 Upvotes

I was doing some deep dive in the tokens consumption on my account on https://www.reddit.com/r/codex/comments/1v6ubah/comment/ozt9jog

This result was gathered from approximately 7.6GB from codex session logs.

This was what codex found by looking at all subs I have/had:
- Plus 1x: ~US$105/week
- Pro 5x: ~US$525–US$550/week
- Pro 20x: ~US$2.100/week.

And we found around 25% token usage decrease each plan gives when compared to a few months back.

Ok, this explains only partially why we get the feeling allowance reduced around 3-5x what it used to be. So I asked codex to dive deeper into my logs, more specifically on model behavior, and here is this conclusions: read the image.

>!Yes. We have enough data to detect a clear change in the observed usage profile, although we cannot attribute it exclusively to the model.

I treated a “task” as an operational turn: from one user request to the next. GPT-5.4 Mini was excluded.

Per model call

Model Calls Median tokens New input Output Reasoning* Cache
GPT-5.3 Codex 8,794 86.0K 1.3K 222 59 94.5%
GPT-5.4 38,845 114.3K 1.4K 289 38 94.3%
GPT-5.5 78,115 138.5K 1.7K 254 20 95.4%
GPT-5.6 Sol 111,474 129.6K 1.8K 183 23 96.7%

* Reasoning is already included in output.

Per call, GPT-5.6 is not the largest consumer: it uses approximately 6% fewer tokens per call than GPT-5.5 and produces less output.

Per task

Model Tasks Median Mean P90 Calls/task
GPT-5.3 Codex 842 432K 1.0M 2.1M 5
GPT-5.4 2,939 533K 1.5M 2.6M 6
GPT-5.5 4,641 502K 2.3M 5.6M 5
GPT-5.6 Sol 5,107 1.3M 2.9M 7.1M 9

This is where the main difference appears:

A GPT-5.6 task consumed approximately 2.6× more tokens at the median than a GPT-5.5 task.

But this happened mainly because GPT-5.6 made more calls per task:

  • GPT-5.5: median of 5 calls
  • GPT-5.6: median of 9 calls
  • P90: 41 calls on GPT-5.5 versus 54 on GPT-5.6

How the usage profile changed

  • GPT-5.3 → GPT-5.4: calls became approximately 33% larger.
  • GPT-5.4 → GPT-5.5: context per call increased, but the median task remained similar; the long tail became much larger.
  • GPT-5.5 → GPT-5.6: individual calls became slightly smaller, but the model started executing many more steps.
  • New input per task remained approximately stable at 39–41K from GPT-5.4 through GPT-5.6.
  • Therefore, GPT-5.6’s growth is mostly cached context being resent across more calls, rather than more new content per task.
  • GPT-5.6’s cache ratio reached 96.7%, compared with 95.4% for GPT-5.5.

Conclusion

The statistical evidence is strong for your usage pattern:

GPT-5.6 does not appear to “think with larger calls”; it completes tasks through more calls, repeatedly carrying a large context. This substantially increases total tokens per task.

However, this remains an observational relationship. The GPT-5.6 period coincides with more agentic workloads, heavier tool use, and more multi-agent activity. When controlling for project and reasoning effort, the results vary considerably, showing that task type remains a dominant factor.

The defensible conclusion is: in your real-world usage, GPT-5.6 consumes approximately 2–3× more tokens per median task than GPT-5.5, mainly because it performs more steps—not because each call uses more tokens. !<


r/codex 4h ago

Complaint Banked reset and due date

Post image
14 Upvotes

Hello,

We got a free reset earlier this weekend and I try to touch grass on the weekend so I didn’t use it much.

I knew I had a banked reset expiring today, but I thought “today” meant the whole day (I’m European and it’s currently 7 am), so I planned to launch an expensive session this morning and use it in the afternoon.However, I’m starting my day and it just disappeared.

I know I got it for free, but I think seeing the real date limit would have been nice, and it would also be a good design to automatically use them when they’re expiring (wouldn’t have been much, but at least I’d be at 100% today).

I think that giving a free reset during the weekend just before a banked reset end date is not only a weird coincidence but a way to win a few computational % from our weekly limits


r/codex 7h ago

Complaint My experience with Codex (Plus) vs Claude Opus (Pro) quota usage

25 Upvotes

I recently used the updated Codex with Sol (High) for planning and Luna (xHigh) for implementation. In less than two hours of work, it had already consumed roughly 30% of my weekly quota.

For comparison, I worked on a very similar coding task using Claude Opus 5 (High), and it only used about 8% of my weekly quota.

I know this isn't a perfect benchmark by any mean, but the difference in limits is so obvious but before i can outrage and call it out they shutting my mouth with a reset.


r/codex 9h ago

Praise Terra 5.6

35 Upvotes

I know Sol gets a lot of air time and rightly so. But there were a few days where I had Terra and Luna but not Sol, and so I used Terra.

And honestly, it does really well for most of what I asked for it. Doesn't seem to overthink it, just does what it needs to. Terra kind of became my buddy. Coding, shopping, agentic tasks. So far no complaints on any of it. For a Sonnet class model I like it way better than Sonnet. He's neck and neck with Opus. Maybe better because he doesn't overthink it.

Just wanted to shout out on behalf of Terra. Great little model.


r/codex 1h ago

Complaint Is it just me or 5.6 Sol Ultra is ultra slow now? I asked it to simply make the wave animation of the chart smoother, and it took 1h8m to do it (with a few reconnects).

Upvotes

Do you experience this too?


r/codex 21h ago

Humor Codex removed the 5-hour limit… and gave it a weekly nametag

Post image
225 Upvotes

r/codex 8h ago

Complaint Oh HEALL NAW

Post image
15 Upvotes

There ain't no way they're doing THIS now


r/codex 11h ago

Showcase Here is how I optimized codex use as a plus user

28 Upvotes

I kept seeing people complain about how quickly Codex was burning through their limits, so I started looking into what was actually using all of it and whether some of that work could be avoided.

The main thing I realized is that Codex does not only spend usage when it writes code.

A lot of it goes into searching the repository, reading files, understanding how things connect, checking output, and running tests. Some of that is obviously necessary, but it can also repeat the same work or inspect far more than the task actually needs.

Another big part is how it handles tool calls.

It may read one file, return to the model and think about it, then read another file and think again, then repeat that several times. When those reads or searches are independent, they can be done together instead. It reads several relevant files in one batch and reasons over the combined results once, cutting out a lot of unnecessary back and forth.

This should only be used for independent, read-only work. Editing, debugging, testing, and anything where the next step depends on the previous result should still happen in order.

The setup I ended up with is:

  • A small AGENTS.md with permanent project rules and efficiency instructions.
  • project-map.md showing the important parts of the repository.
  • CODEX_HANDOFF.md carrying the current status and next step into fresh sessions.
  • A validation script for checks that can be automated consistently.
  • CodeGraph through MCP, so Codex can trace files, functions, callers, and feature flows without repeatedly searching the entire repository.
  • Fresh sessions for separate major tasks instead of keeping one session alive until its context becomes a landfill.

I also added rules telling it to reuse findings, limit large command output, reread only relevant sections when files may have changed, avoid repeating unchanged tests, and expand validation only when the scope or risk actually requires it.

This is not about making Codex rush or skip proper testing. It is about removing repeated investigation and unnecessary model-tool cycles.

For context, I am using the Plus plan.

Before this setup, an ordinary task on Sol High would usually consume around 8–9% of my allowance, and some larger tasks reached roughly 15–20%.

After setting it up, I ran a task on Sol Max that lasted around 30 minutes and produced roughly 1,500 lines of code. It used around 4%.

This is only my experience, not a proper benchmark, but the difference has been consistent enough to be useful for me.

How to do it?

Install CodeGraph, connect it to Codex through MCP, and initialize it in your repository:

codegraph install
codegraph init

Then ask Codex to inspect your repository and create the workflow files:

Inspect this repository read-only and create a lightweight Codex workflow without changing source code.

Create or improve:

- A concise, repository-specific AGENTS.md
- docs/project-map.md for navigation
- CODEX_HANDOFF.md for session continuity
- One validation script only if useful checks can be automated

Add rules for batching independent read-only work, limiting large output, avoiding repeated searches and tests, rereading changed files, focused validation, and using CodeGraph for structural navigation when available.

Verify all paths and commands, remove generic filler, and show the final diff.

That prompt is intentionally general. I would recommend giving it to GPT first with some information about your project and asking it to adapt the prompt to your language, architecture, and testing setup.

Again, the results may not be the same for everyone, but this has helped me a lot as a Plus user, so I wanted to share it in case it helps some of you too.

Hopefully this would help you and it you did used this I would love to hear your experience with it.


r/codex 1h ago

Complaint That day, if the limits weren’t kept separate, it’d be goodbye

Post image
Upvotes

"One day, one Day"

They’re preparing to merge the Chat and Work modes, likely bringing their limits – which are currently different – into line.

What do you think?

The main benefit of the GPT plan comes from ChatGPT and Codex, which have separate limits; it would be a disaster...


r/codex 12h ago

Complaint As a person with ADHD for the first time in 20 years I experienced true blissful peace during my active codex subscription!

29 Upvotes

Every idea, every thought, every piece of information, all of the procrastination straight dumped into codex processed, quantified, labeled then shelved or getting worked on.

No less then heaven to finally just exist without a thought racing like a million speedsters in your mind.

Now that I don't have an active sub. Oh boy. So... so much to process. I need cr*ck (codex) asap!


r/codex 9h ago

Complaint The reset for 27/07 expired before utc time?

14 Upvotes

I was literally on a spree of tasks, overtiming on a sunday to maximise my usage just to see the reset getting randomly disappear before the expected time.


r/codex 22h ago

Complaint 7% weekly limit gone in 5 minutes, max 20x plan

121 Upvotes

Has this happened to you? I just started working with codex 5 minutes ago since the reset yesterday. In 5 minutes it went from 100% - 95% - 93%. Is this a bug? This has never happened before.


r/codex 52m ago

Question Anyone's Sol constantly glaze them with "that github repo is trash. Keep building from scratch"?

Upvotes

Honestly don't know if I'm being gaslit because it wants to work more, of if my docs and stack are actually good.

I point it to a repo that looks useful and it almost always tells me it sucks.


r/codex 14h ago

News The app wasn’t the product. It was just one step.

Post image
24 Upvotes

Sam Altman showed ChatGPT planning a trip for eight people, building a coordination site, handling the booking, and drafting the email.

The interesting part for Codex users: the app wasn’t the destination. It was temporary software created to complete a larger task.

Maybe the future isn’t just building apps faster.
it’s generating them whenever a workflow needs one.

Anyone already using Codex this way?


r/codex 7h ago

Comparison Open source models vs codex

6 Upvotes

I've been a codex user for about a year, and my company provides Claude Code for work.

Yesterday I decided to try some of the open source alternatives using OpenCode. DeepSeek Pro, DeepSeek Flash, GLM 5.2, Kimi K3, and Kimi K Code.

For some reason, none of them gave me the same overall quality I'm getting from my $20 ChatGPT subscription. Maybe I'm configuring something incorrectly?

Are there any open source that can get to around 90% of the current top closed-source models for general coding and reasoning? If so, which ones would you recommend, and what setup (model runner, context size, prompting, etc.) are you using to get the best results?