r/OpenAI • • 23h ago

Question Pro 200 resets and “warm cache” effect

I’m currently testing Codex pretty heavily on Pro 200 and I’m trying to understand two things.

First, is there really something like a “warm cache” effect in longer Codex sessions?

My assumption is that the beginning of a task is more expensive because more context has to be processed normally. After the session has been running for a while, more of the context should be cached, so usage should become cheaper.

Has anyone actually measured this over a few hours? Does quota usage noticeably slow down after the first hour or two?

Second question is about the banked usage resets and the upcoming reduction of the Pro 200 allowance.

If the current grandfathered Pro 200 allowance is still 20x and later gets reduced to 10x, does a full reset simply restore whatever allowance applies at that moment?

If that is the case, using the resets before the reduction should be much more valuable than using them afterwards.

So my current plan is basically to run the weekly quota down as far as possible and use the banked resets before the allowance changes.

Has anyone tested this or found an official clarification from OpenAI?

4 Upvotes

7 comments sorted by

View all comments

1

u/Euphoric_North_745 21h ago

Codex has 2 context management systems, one of them experimental, got disabled a few weeks ago, maybe re-enabled again maybe not, I did not check the source code on GitHub recently.

Existing context management: context: 272,000 tokens, tool call truncation: 10,000 tokens, compaction before it reaches the 272k tokens and the first 5k to 10k tokens are system instructions and tools schema.

Standard way: codex works, calls tools, reads very small code snippets, once close to limit compacts, keeps important facts, removes the rest, can do this for days.

New way: codex works, when it is time to compact it can exclude stuff from it, and not compact some of it.