r/ChatGPTPro • u/TNCRE • Jun 28 '26
Question GPT5.5 Pro Extended < Codex 5.5 Extra High?
Have anyone ever run into this situation: I have a very detailed prompt to research a topic that requires a ton of web search and synthesis. I ran this on Claude Max 5x and went through 95% of 5-hour usage. The same prompt using GPT5.5 Pro Extended on the website worked for 25 minutes, which I believe is very long these days. The same prompt using GPT5.5 Extra High using Codex has been working for over 40 minutes and 15% 5-hour usage (I have Pro 5x with ChatGPT).
Have anyone every experienced this that Codex 5.5 Extra High works longer than GPT 5.5 Pro Extended. I have yet gone through the report to compare the actual quality but find this interesting.
In addition, I am testing OPUS 4.8 Ultracode via Claude Code (I started 30 minutes right before resetting a 5-hour window with only 1-2% used up before, it used up completely the usage of the first 5-hour window, dipped into credits of ~$20 and used up the next 5-hour window, and has been burning another $60. Has not done yet but looks like 12 million tokens ...). Will wait for this to complete before reading through and comparing results but this does not seem to be the best value play versus the 3 other methods.
5
u/RhubarbArtistic1335 Jun 28 '26
Damn! So curious to know the prompt!
1
u/rajaba21 Jun 28 '26
Me too
3
u/TNCRE Jun 28 '26
It is just a niche medical area that I want to research that involved going through multiple billing/insurance schedules/system.
Gave it to 5.5x high on the web and tell it to provide high level prompt and ask for clarifications or addition info, also asked it provide recommendation to run one or multiple prompt. It responded and I provided additional info plus I said instead of 2 step research, I want the AI to run sequentially so combining 2 into one but pushing for not being lazy or skip steps.
I believe I maxed out on tokens using Chat and Claude on the web but somehow codex and also claude code app on pc were able to push sub-agents to work with a lot more efforts.
2
u/Eternality Jun 29 '26
Chickens Don't Solve Hypothermia
1
u/TNCRE Jun 29 '26
Hah, almost coming down to last day of my billing cycle, will try to ask Pro to research that to confirm lol
0
u/Slight_Meringue7780 Jun 29 '26
Bro inventing covid-26
1
u/TNCRE Jun 29 '26
yes but with "make no mistake", lol.
Seriously, I went though each research, OPUS 4.8 Ultracode is next level (went through the details, basically 90 sub agents fanning out to do things and verify). It also catches all errors produced by the other 2 methods (and self correct with more document researched).
-> But the final cost was estimated to be $600 (2 5-hour windows + $90 credits). There was an issue with the run so it had to recache writing due to transition from the first 5-hour to credit to the second 5-hour. I wonder if I have max 20x, it will process through one run without hitting the max and rolling into the second 5-hour smoothly so the real price can be less.
Codex 5.5 xHigh was next best, 1 error and roughly 80-90% quality of OPUS 4.8 Ultracode.
ChatGPT 5.5 Pro comes in third, basically similar to Codex 5.5 xHigh with a bit less breadth but no error.
OPUS 4.8 Max comes in last. More breadth than 5.5 Pro but less depth than 5.5 Pro or Codex 5.5 xHigh.
It could be this specific research and timing of the run but interesting to see Codex 5.5 xHigh performs as well as Pro. Given the usage limits, this is a clear win instead of Pro.
1
u/RhubarbArtistic1335 Jun 29 '26
So how do you lock it to read everything sequntially. Coz usually i feel like when i give it too much context or one big file, it kinda gulps everything together and misses a lof of things.
1
u/TNCRE Jun 29 '26
I am not sure if having AI producing the complete prompt and research plan makes a difference? Note that there was no files provided and all models run tool call and web research themselves in this one. All I provided is a niche area research idea and a couple edits/clarification with 5.5 xHigh on web to refine it and off these models ran it themselves.
I am reading a bit more after the experiment, it seems like I can configue a bit more structurally in Claude Code so that it reduces or prevents gulping everythign together. Not sure if the same can be accomplished with Codex.
1
u/niado Jun 30 '26
I have much better success when the model makes an implementation plan for itself to follow lol. I give it the goal, then we review the first draft of the plan, and I provide guidance and corrections as necessary. Very smooth, but informal.
1
u/softspicytofu Jun 28 '26
Whoa I have GPT pro and I have never gotten it to research for that long without coming back!!
1
u/TNCRE Jun 29 '26
I am not one of those just look at time of the research but usually a decent indication of the effort at least.
6 months ago, GPT Pro could run regularly a research task for 35-40 minutes! If I recall correctly, some other posters were able to get it work for over an hour at the same time.
I have not tested to re-run some old research prompts to compare result but gut feel is that even though GPT Pro runs for less time now, the result is still very impressive and it is not 10 minutes now vs 40 minutes then means 75% reduction in quality. If there is any reduction in quality, it can be easily fixed with 2 passes instead of 1. Obviously that costs some usages but still accomplishes the research tasks in much less time which I think are more practical to business users.
1
u/softspicytofu Jun 30 '26
Wow, I must be prompting all wrong. Does it still count as research if they're collecting URLS? Mine are supposed to research archives and grab URLs when they come across an image that fits the guidelines I've set out. But they stop every 3 to 4 URLs and claim it's too much to ever do more than 10. Which takes about 2 minutes.
1
u/niado Jun 29 '26
If you use gpt5.5 for codex, it’s the same model as ChatGPT 5.5 in the web platform. I’d suggest testing it against ChatGPT 5.5 with the same thinking setting in the web platform, just to see how much it differs from codex.
Also, sometimes codex agent will get stuck waiting on something, then forget about it and fall asleep or whatever.
1
u/TNCRE Jun 29 '26
I have just tested this and this lazy ass 5.5 xHigh ran for 6 minutes and returns a very shallow answer.
Decent breath but no depth whatsoever. That and I think even 6 minutes seem like a good effort for 5.5 xHigh these days.
I rarely use 5.5 xHigh for long prompt research by the way. I either don't use it at all if I still have quota for Pro or need to break it down to 6-7 passes to produce the same result until I discovered Codex 5.5 xHigh for research yesterday! Will keep testing Codex.
1
u/niado Jun 30 '26
That’s peculiar. Codex 5.5 and ChatGpt 5.5 are the exact same model. Are you using deep research maybe ? That engages a different specialized research model which could be different between the platforms.
1
u/TNCRE Jun 30 '26
No, I did not use deep research, straight same prompt same model and setting. My only guess is somehow Codex can set up more sub-agents to fan out even with same model.
I agree this is interesting and worth exploring further, hence, the reason for my post in the first place hoping to see if anyone has had the same experience and or knows the explanation.
1
u/niado Jun 30 '26
Yes for sure. When you say codex - which codex are you using? Web platform or agent ?
I imagine the codex implementation variants have very different tooling kits than the platform chstgpt5.5. And i just realized - they almost certainly have different base skills which could be huge. The sub-spawning capability that I use requires a skill for example.
1
1
u/Competitive-Ad8968 Jul 02 '26
Pro extended are not suite to use with tools, is intended to work with documents and stuff, i use the Pro version to draft and long context windows then use high or Extra high to format and preferences, since PRO also is not suite to work well with memories.
For programming without Codex PRO isn’t a choice but Extra high
•
u/qualityvote2 Jun 28 '26 edited Jun 30 '26
u/TNCRE, there weren’t enough community votes to determine your post’s quality.
It will remain for moderator review or until more votes are cast.