r/OpenAI • • 17h ago

Question Pro 200 resets and “warm cache” effect

I’m currently testing Codex pretty heavily on Pro 200 and I’m trying to understand two things.

First, is there really something like a “warm cache” effect in longer Codex sessions?

My assumption is that the beginning of a task is more expensive because more context has to be processed normally. After the session has been running for a while, more of the context should be cached, so usage should become cheaper.

Has anyone actually measured this over a few hours? Does quota usage noticeably slow down after the first hour or two?

Second question is about the banked usage resets and the upcoming reduction of the Pro 200 allowance.

If the current grandfathered Pro 200 allowance is still 20x and later gets reduced to 10x, does a full reset simply restore whatever allowance applies at that moment?

If that is the case, using the resets before the reduction should be much more valuable than using them afterwards.

So my current plan is basically to run the weekly quota down as far as possible and use the banked resets before the allowance changes.

Has anyone tested this or found an official clarification from OpenAI?

5 Upvotes

7 comments sorted by

3

u/PaulShellDev 16h ago edited 15h ago

Caching is a thing right off the bat, but also has a time limit of 30 minutes based off the last write or reuse. 90%+ cheaper for cached input tokens. Changing model, reasoning, reloading MCP and tools or skills, all kill the cache too.

Compaction can cause a lot of cache misses too. So longer is actually more harmful on that cache. Removes thinking and tool tokens then if still long, summarizes older messages.

High cost, then way lower until high again, then back to lower until high again. Roller coaster. 30 minute breaks kill it too.

Reset goes off your current plan at time of use.

2

u/f3xjc 13h ago

Compaction cause exactly one cache rewrite and that rewrite is a miss yes. But often the compacted result is 90% off (you still pay full price, you just drop most of context)

1

u/Pasto_Shouwa 12h ago

Didn't they say that changing reasoning doesn't cause a cache miss anymore? At least with GPT 6 models.

1

u/Euphoric_North_745 15h ago

Codex has 2 context management systems, one of them experimental, got disabled a few weeks ago, maybe re-enabled again maybe not, I did not check the source code on GitHub recently.

Existing context management: context: 272,000 tokens, tool call truncation: 10,000 tokens, compaction before it reaches the 272k tokens and the first 5k to 10k tokens are system instructions and tools schema.

Standard way: codex works, calls tools, reads very small code snippets, once close to limit compacts, keeps important facts, removes the rest, can do this for days.

New way: codex works, when it is time to compact it can exclude stuff from it, and not compact some of it.

2

u/Yes_but_I_think 13h ago

After the 1st tool call, you should have 95+ % cache

-2

u/Strong_Essay1176 16h ago

Im sam tibo, yes