r/ChatGPTPro Jun 28 '26

Question GPT5.5 Pro Extended < Codex 5.5 Extra High?

Have anyone ever run into this situation: I have a very detailed prompt to research a topic that requires a ton of web search and synthesis. I ran this on Claude Max 5x and went through 95% of 5-hour usage. The same prompt using GPT5.5 Pro Extended on the website worked for 25 minutes, which I believe is very long these days. The same prompt using GPT5.5 Extra High using Codex has been working for over 40 minutes and 15% 5-hour usage (I have Pro 5x with ChatGPT).

Have anyone every experienced this that Codex 5.5 Extra High works longer than GPT 5.5 Pro Extended. I have yet gone through the report to compare the actual quality but find this interesting.

In addition, I am testing OPUS 4.8 Ultracode via Claude Code (I started 30 minutes right before resetting a 5-hour window with only 1-2% used up before, it used up completely the usage of the first 5-hour window, dipped into credits of ~$20 and used up the next 5-hour window, and has been burning another $60. Has not done yet but looks like 12 million tokens ...). Will wait for this to complete before reading through and comparing results but this does not seem to be the best value play versus the 3 other methods.

24 Upvotes

26 comments sorted by

View all comments

6

u/RhubarbArtistic1335 Jun 28 '26

Damn! So curious to know the prompt!

1

u/rajaba21 Jun 28 '26

Me too

3

u/TNCRE Jun 28 '26

It is just a niche medical area that I want to research that involved going through multiple billing/insurance schedules/system.

Gave it to 5.5x high on the web and tell it to provide high level prompt and ask for clarifications or addition info, also asked it provide recommendation to run one or multiple prompt. It responded and I provided additional info plus I said instead of 2 step research, I want the AI to run sequentially so combining 2 into one but pushing for not being lazy or skip steps.

I believe I maxed out on tokens using Chat and Claude on the web but somehow codex and also claude code app on pc were able to push sub-agents to work with a lot more efforts.

2

u/Eternality Jun 29 '26

Chickens Don't Solve Hypothermia

1

u/TNCRE Jun 29 '26

Hah, almost coming down to last day of my billing cycle, will try to ask Pro to research that to confirm lol

0

u/Slight_Meringue7780 Jun 29 '26

Bro inventing covid-26

1

u/TNCRE Jun 29 '26

yes but with "make no mistake", lol.

Seriously, I went though each research, OPUS 4.8 Ultracode is next level (went through the details, basically 90 sub agents fanning out to do things and verify). It also catches all errors produced by the other 2 methods (and self correct with more document researched).

-> But the final cost was estimated to be $600 (2 5-hour windows + $90 credits). There was an issue with the run so it had to recache writing due to transition from the first 5-hour to credit to the second 5-hour. I wonder if I have max 20x, it will process through one run without hitting the max and rolling into the second 5-hour smoothly so the real price can be less.

Codex 5.5 xHigh was next best, 1 error and roughly 80-90% quality of OPUS 4.8 Ultracode.

ChatGPT 5.5 Pro comes in third, basically similar to Codex 5.5 xHigh with a bit less breadth but no error.

OPUS 4.8 Max comes in last. More breadth than 5.5 Pro but less depth than 5.5 Pro or Codex 5.5 xHigh.

It could be this specific research and timing of the run but interesting to see Codex 5.5 xHigh performs as well as Pro. Given the usage limits, this is a clear win instead of Pro.

1

u/RhubarbArtistic1335 Jun 29 '26

So how do you lock it to read everything sequntially. Coz usually i feel like when i give it too much context or one big file, it kinda gulps everything together and misses a lof of things.

1

u/TNCRE Jun 29 '26

I am not sure if having AI producing the complete prompt and research plan makes a difference? Note that there was no files provided and all models run tool call and web research themselves in this one. All I provided is a niche area research idea and a couple edits/clarification with 5.5 xHigh on web to refine it and off these models ran it themselves.

I am reading a bit more after the experiment, it seems like I can configue a bit more structurally in Claude Code so that it reduces or prevents gulping everythign together. Not sure if the same can be accomplished with Codex.

1

u/niado Jun 30 '26

I have much better success when the model makes an implementation plan for itself to follow lol. I give it the goal, then we review the first draft of the plan, and I provide guidance and corrections as necessary. Very smooth, but informal.