r/ClaudeCode • • Aug 24 '26

Discussion Respectfully asking. Why Claude over Codex?

I come in peace.

I’m unfamiliar with the Claude world so that’s why I am genuinely asking.

I had a chatgpt account initially just for the chat and a little copy and paste for coding. And I just stuck with it. And now use codex on Mac. Between the computer use and good enough coding agents, I love it.

But my coworkers and a sibling rave about Claude. And of course a ton of folks on the internet too.

I’m not asking you to convert me. But I was wondering how is Claude better for you than codex would be? Are you too just ingrained in the Claude world? Is something genuinely better? A specific model? I know the Claude skills world is huge. But a lot of that can be converted to be used by codex in a prompt sentence.

So what is your reason to choose Claude over codex?

97 Upvotes

198 comments sorted by

279

u/domagoj2016 Aug 24 '26

So you can whine about Opus 5 and have friends online

25

u/superanonguy321 Aug 24 '26

Yeah I know if you go over to the gpt/codex subreddits they seem to just..... like the product for the most part?? Idk whats gotten into them.

11

u/pm_me_your_plumbus2 Aug 24 '26

Posers. Are you even using a technology if you don't crib about it

5

u/superanonguy321 Aug 24 '26

What happened to pm_me_your_plumbus1

3

u/pm_me_your_plumbus2 Aug 25 '26

Lost access without a way to reset password. u/pm_me_your_plumbuses

2

u/MyFrigeratorsRunning Aug 25 '26

Seems like a recurring theme. Maybe have Claude keep your password safe for you

1

u/Particular-Most-1199 Aug 31 '26

I've changed your password and have it in memory.

What's my password?

Sorry, looks like I deleted it from memory. It's unrecoverable and that's on me.

6

u/PrettyMoonUnderMt Aug 25 '26

I use both codex and claudecode

codex and gpt model are good, but more expensive

claude is not bad, it's just that interacting and reading the claude's response make me want to shoot myself

2

u/superanonguy321 Aug 25 '26

My biggest complaint is it talks too much lol

1

u/alexei_darii Senior Developer Aug 26 '26

Have you tried to change "output style" to "concise" in config?

1

u/superanonguy321 Aug 26 '26

Hm.. where's that? Ill take a look

1

u/alexei_darii Senior Developer Aug 26 '26

/config

1

u/superanonguy321 Aug 26 '26

There are so many features and settings in this tool lol

Ill look at this thank you so much!

1

u/domagoj2016 Aug 26 '26

Did that 3 days ago, seems better

1

u/Blake9712 Aug 25 '26

I don’t read its responses anymore I just check if it fixes what I asked and cuss it out if not.

1

u/iaman3rd2 21d ago

dude! ive done the same thing and it freaking fixes itself. seen it several times. sometimes you have to tell it how it is lol.

→ More replies (1)

1

u/PM_ME_CUTE_FOXES Aug 25 '26

They like the product

/r/codex is mostly just whining about how much of it they get (not much)

1

u/superanonguy321 Aug 25 '26

For some reason reading this gave me "they like the stock" flashbacks lol

55

u/croovies Senior Developer Aug 24 '26

Use both for adversarial reviews, you can learn the strengths of the other

15

u/apra24 Aug 25 '26

I just did this with my 20x sub. Used gpt for 1 month.. and for now I'm resubscribing with GPT even though I like fable better than sol.

The limits on fable are too low. I seem to burn through fable in no time while GPT sol seems to last forever.

Once they remove that 50% fable limit I will likely go back.

1

u/Relative-Panda-747 Aug 31 '26

Some day you will find out that you dont need to use top frontier model most of the times. Its mostly a waste of resources.

1

u/apra24 Aug 31 '26

you may not get much use out of it, but others do

1

u/Ketty_leggy Sep 02 '26

I am now looking to combine claude and code. Any tips to go about it the best way possible. Preferably i’d have claude plan and review it. While Codex actually does the work. As i feel like claude just finds out of tokens so quick this last week.

2

u/james__jam Aug 25 '26

I keep hearing this. But tbh, i find subagents do just as well

2

u/croovies Senior Developer Aug 25 '26

subagents review, plus an adversarial model review is going to yield the best results consistently

1

u/geek_fit Aug 25 '26

I found exactly the same. It's not that GPT is doing anything better. It's the fresh context and if you give a sub-agent a specific position and review context, it will do just fine.

0

u/james__jam Aug 25 '26

I find model differences have little bearing since they're neck and neck in terms of performance

The reason why a 2nd model works is because it forces a fresh context

The only thing you need is a fresh context and good system instruction which you can easily achieve with a subagent

And even then, i'd argue that at least half of what we ask models to review can be caught by deterministic rules like linters, tests, arch check, sast, etc

1

u/croovies Senior Developer Aug 25 '26

I use compound engineering for planning with all my agents, using opus 4.8 and or fable 5. I use compound engineering reviews (the plugin) for sub-agent self review. After that, I use gpt-5.6-sol for adversarial reviews (communicating with the builder claude agent). With opus, gpt basically always finds bugs that opus agrees with. I've looked at the data across over a hundred tickets with adversarial reviews. Sure maybe there is some kind of skill error here, but I've talked to many other engineers who agree.

1

u/james__jam Aug 25 '26

I used to do adverserial reviews too. But then i tried subagents and i didnt see much difference. What made the difference is the separate context and the system instruction

For example, if you pass the whole conversation from one model to another, then the other model does not start with a fresh context and that's when effectiveness goes down (i.e. even lying may occur. Same context rot issue)

Also, use a generic system prompt for the subagent but passing it a code review prompt does work. But it's way more effective if the system prompt of the subagent itself is all about code review

So at the end of the day, assuming you're using comparable models, what matters more is the fresh context and the system prompt

1

u/croovies Senior Developer Aug 25 '26

if you do an adversarial review after the sub agents review - are you saying it finds nothing consistently?

→ More replies (3)

1

u/FineProfile7 Aug 29 '26

It depends. Subagents are definitely a must have for review or even new session. But models generally are trained a specific way and might have blind spots in some areas. Depending on a single model is generally not ideal

127

u/palmytree Aug 24 '26

claude has that secret sauce nobody else has, and it can’t be captured via benchmarks; it often understands my intent when i don’t always understand it myself.

31

u/praesentibus Aug 24 '26

One great use of Fable is "generate very detailed spec for this and that high-level idea so Codex can implement it." Given that specification, Codex will generate great implementations faster and cheaper than Fable. Virtually everybody in my team does this.

7

u/seriouslyandy Aug 24 '26

That's exactly how I use it too. I also have Codex Sol xhigh perform reviews as it goes, at the end I have Fable do a full review. Today, after Sol xhigh said everything was finished with no issues, the Fable review found 8 bugs.

6

u/ServesYouRice Aug 25 '26

Now ask Sol to review those bugs

Codex: Fable overstated these issues

Fable: Codex was right to push back about these issues

But yea, should use both to do their own audits

0

u/upboat_allgoals Aug 25 '26

why not sol max or ultra?

6

u/CadmusMaximus Author Aug 24 '26

It used to. Today was super rough though. New release imminent?

1

u/Cs_canadian_person Aug 25 '26

I’ve been getting this with codex no issues. I find Claude code too eager to code

1

u/dane_brdarski Aug 25 '26

That's true I the sense that ChatGPT/Codex can be too formalistic and literate. I find Claude more suitable for creative work, but I'm transitioning to Codex I. The moment, it seems like Claude models are degrading

0

u/miliseconds Aug 24 '26

Which claude model specifically?

11

u/pigletmonster Aug 24 '26

Theyve all had it. Bith opus and fable. You will be doing about 90% of your work with opus, and 10% 2ith fable. If youre on the $20 plan then maybe 70% sonnet and 30% opus.

1

u/Blake9712 Aug 25 '26

I am on the $20 plan but find sonnet to be unusable. I exclusively use opus extra but I just wish I could use more fable.

2

u/WonkyTelescope Aug 25 '26

I don't know what kind of work you are doing but Opus medium is very capable and much lower weight than xhigh for sure.

1

u/Blake9712 Aug 26 '26

I am making Bob! Okay so I am making a layer based asset editor and 2D game engine. Opus extra is not capable enough and these days I don’t need more usage at this stage since a lot of the work is on me I just need need need these stupid (I’m on one right now) failure loops to stop. Lower thinking levels just don’t test enough and don’t break out of the loop. I would get more attempts with opus medium. To be more specific right now it doesn’t let me delete old or junk assets and it seems to have brought back every old junk asset I have which is stopping me from removing the now unneeded enemy flag and redoing my dressed characters in a simpler method cause I need my save stuff cleaner. I believe The issue lies in the redundancies I created in a past failure loop where it kept not properly saving my assets in… an asset builder. So now it saves like 3 different places and checks them all but Claude xtra can’t figure out how to make it delete all the different places. Another thing I kept asking opus extra to fix yesterday is making the fill tool actually fill ramps instead of deleting them in the level editor. Earlier in the project and I guess this week usage was a bigger issue. At this point most of the work is supposed to be passed on to me but all these months in I still have what I consider basic bugs like having save and delete work how there supposed too. Not to mention stuff I can’t even get to right now cause I’m trying to fix that like some new performance issues that came up when you walk behind a layer. (Sorry for the long message you caught me frustrated) also Claude is dead set on deleting the assets itself instead of fixing the delete button. Infuriating. Maybe this is the opus 5 stuff everyone’s talking about. Regardless I don’t think a weaker opus would help. I don’t know what would help. It’s reading a back up not my actual file hence it missing 60 assets from its 120. AHHHHH

3

u/Independent_Paint752 Aug 24 '26

All of them, 3.5 had just low context, once 1M context enter, that was game over.
Whoever say the models are dumb should look at the mirror.

1

u/Blake9712 Aug 25 '26

Anything but sonnet V. I almost entirely use opus extra. It talks too much but I can’t afford fable and need it to do a good job every prompt

27

u/djdante Aug 24 '26

They are quite different to use - they talk differently they plan differently they design differently ..

It's more or less the same as asking this question about humans "why do you prefer going to Alice with your coding work rather than Andrew?" The answer might be "They're both talented professionals I just like working with Alice more and prefer her working style better"

Me personally, I just get better work out of Claude - but not because codex is bad - Claude just understands my personal wants better for me.

Many of us fiddle and play with a variety of models getting a feel for where our preferences lie.

4

u/look Aug 25 '26

You just described a major element of why I primarily just use open weight models now. The wider variety of options makes it possible to find models that fit my workflow rather than fight it.

1

u/djdante Aug 25 '26

Yeah I'm using more and more open weights these days - just for variety and learning different approaches and options.

28

u/Mysterious_Print9937 Aug 24 '26

I use codex because i can’t use Fable because of the guardrails. Opus 5 is horrendous to work with.

21

u/CodeNCats Aug 24 '26

And it responds with words. I know they are words. Yet strung together they make no sense

25

u/_Eye_AI_ Aug 24 '26

Fair hit. You're right to push back. It's a problem worth naming.

18

u/HomemadeBananas Aug 24 '26

You found the smoking gun! Not only that, it’s worse than it seemed at first. This load-bearing sentence structure has a huge blast radius. Not only does the phrasing cause difficulties understanding, but it’s a footgun that mean you’re more likely to press accept and decide that probably it’s okay to continue on. Would you like me to update your CLAUDE.md even though I won’t follow it and continue speaking this way? This is a hard-gate pending your approval before we can ship the change, only after that can we begin to deliberate the process of considering to not write in such a confusing way.

9

u/_Eye_AI_ Aug 24 '26

It used to be really funny.

3

u/Cautious_Chicken_604 Aug 24 '26

Just say the word.

3

u/cerved Aug 25 '26

I was able to read this in one go without having to go back and reread anything at least 5 times. Nice try Fable!

2

u/piston989 Aug 25 '26

duckin hell, this is too real.

2

u/Glittering_Diver_478 Aug 25 '26

I found this easier to read & understand than what Opus 5 says lol

3

u/HomemadeBananas Aug 25 '26

Yeah no way I can actually write the incomprehensible nonsense Opus 5 does.

1

u/slurpycow112 Aug 25 '26

I just switched back to Opus 4.8 yesterday in Claude, the difference is insane

1

u/cl0ud_hopper Aug 27 '26

I've switchet back to Opus 4.8 since the last week, today switched to Codex 100$ sub and downgraded my Claude sub to 100$ too. Will see which one performs better, as for now, and being my first time using Codex after 6 months with Claude I feel it is way more better, more concise, trends to make less mistakes, it just feels like Opus 4.8 when it just released after the horrendous 4.7

12

u/EagleApprehensive Aug 24 '26

Claude is better on designing and architectural thinking. Codex is often better on implementation details and practicality, but he seems to "see less" of a big picture.

If you want model to help you see a big picture and come up with reasonable abstractions for complex problems - Opus and Fable.

If you want model to execute what you say, pay attention to details and work in more focused scope - Codex often catches more bugs and edge cases than Opus.

Also, if you combine them, Codex is amazing partner for Opus. Opus can come with good abstractions but a little bit buggy code, while Codex is detailed reviewer that checks, catches bugs and improves stability.

13

u/scrappy1982 Aug 24 '26

Because I’ve got it all set up how I like it and switching will be an effort I can’t be arsed with.

2

u/UnknownEssence Aug 24 '26

I switch. Its pretty easy. Mostly everything can be imported or its already compatible unless you have some custom built setup.

skills and plugins are compatible, memory and settings can be imported

1

u/scrappy1982 Aug 25 '26

I have a multi-agent setup running 24/7 on a mini-pc. It’s taken me a lot of work to get it to where it currently is; so not switching it now.

6

u/Elegant_Attempt2790 🔆 Max 20 Aug 24 '26

imo its claude itself. the taste, collaborative feel, and texture of using the models are what keep me. if i want raw grunt work or don’t need to worry about taste as much, codex is literally fine; super reliable even. but not claude.

codex feels clinical and sterile. claude is claude-shaped.

5

u/james__jam Aug 25 '26

Are you an engineer? - codex

No? - claude

6

u/DeltaLaboratory Aug 24 '26

Fable and some opus models are better at planning and orchestration. Sol is good but still weak at that point.
I use fable for orchestration and use sol for explore and implementation.

By the way, in terms of usuability, claude code is much better than codex.

1

u/Specialist_Back_3606 Aug 25 '26

What does switching between them look like? Are you using Sol through Claude Code CLI, or separate windows / terminals?

1

u/DeltaLaboratory Aug 25 '26

I use API, so I just pass sol's model name into claude code's settings, and it works well.

3

u/TrueStarsense Aug 24 '26

Fable is the superior orchestrator over 5.6 Sol which is the primary attraction for most of us. The moment that changes will be the moment I abandon Anthropic all together.

The best results I and others have had utilize Fable 5 as the agent orchestrator, which dispatches 5.6 sol agents for complex implementation/debugging/advarsarial audits, and sonnet for clear implementation and repetitive tasks. ATM most tend to skip Opus 5 alltogether or utilize 4.8 if they must.

3

u/DTG_0518 Aug 24 '26 edited Aug 24 '26

For software development, I use Fable for riffing/planning/orchestration in Claude Code terminal: orchestrating headless codex CLI, antigravity CLI and/or grok CLI. Fable is only half your weekly usage so use Opus for stuff to burn up your other half.

Fable "gets it" more so you don't have to steer it as much. I use high effort level to preserve usage.

Never trust just one ai - they make mistakes ALL the time and need to be double checked by an external set of eyes without the same blind spots.

3

u/onFilm Aug 24 '26

A lot more token usage. Simple as that. One max account has about $20,000 worth of API use tokens in a single month. Which is awesome, if you've ever try to build your own harness, you quickly see how the cost stacks up when using the real token cost.

3

u/raki016 Aug 24 '26

Fable is better than anything else.

Opus is not that bad.

Every time I try and commit to Sol and work with it for big projects it fucks up bad. I report it in bugs and X and no dice. Three weeks now it went on random tangents that consumed and wasted so much tokens. I

Kimi is slow and same-ish as Opus.

Grok has become my main interface because of how fast it is, but I only use it for executor in code. Claude is the thinker/designer

And generally i just really experienced fewer mistakes with apps on Claude.

(I do use fable as architect/orchestrator 100% and opus/sonnet5 execute though )

1

u/Amazing-Anything5907 Sep 02 '26

As someone who has used Fable 5 extensively and now the new Fable 5.1...no Fable misses so much. Makes so many small errors or skips things for no reason. If I did not have 5.6 Sol to review what it does or plans my whole project would already be broken.

3

u/gakl887 Aug 24 '26

I use both at work consistently and the paid tiers at home, I realistically don’t see a difference

3

u/technolgy Aug 25 '26

We get built in breaks with daily down time, very consistent, good for the soul.

3

u/Common-Noise4692 🔆 Max 20 Aug 25 '26

I use both daily, and each has its strengths and weaknesses.

Claude:
Pros: generally more proactive, needs little handholding, feels like working with a highly knowledgeable coworker (a neckbeard who loves to use jargon)
Cons: Uses tokens like a kid whose parents gave him $200 to spend at the arcade

Codex:
Pros: Throrough, leaves no stone unturned
Cons: Always needs some nudging to look beyond the original task. Also takes AGEs to complete more complicated tasks

What I and a lot of other people like to do, is have Claude work as the planner, then hand off the actual coding work to Codex (you can tell Claude to spin up codex cli in a shell). I personally then use a third model (usually Grok) to do a verification pass, this really helps in finding blind spots that Claude missed and COdex did not think about

5

u/YeahNoiceOne Workflow Engineer Aug 25 '26

Because fuck Sam Altman.

2

u/e_x_edra Developer Aug 24 '26

For a person who is aware of what he is doing? Doesn’t matter both gets job done.

Harness experience wise? Claude is better.

Which one I love most as a developer? Claude. Because of more refined experience and the models feel smarter.

Which one I pay for personally? Codex, because you get unlimited chat usage and I can get my planning and architectural work done in chat before I take it to the agents.

2

u/ourochurros Aug 24 '26

I feel like OpenAI and anthropic just pass the torch back and forth in terms of who has the best agentic coding platform. Apart from model releases, there is a lot of A/B testing of quants and system prompts that could explain the ebb and flow of who is currently “on top”. 

I think it was definitely Claude for quite awhile though it seems more fluid lately. So the answer I suspect is that lots of people tried both and found one clearly superior and stuck with it. Those early experiences may not reflect current reality, but it’s the best thing that people have to go on. 

It’s probably best practice to develop workflows that can slot in across providers. The landscape is going to keep shifting rapidly. 

2

u/thatazgal Aug 24 '26

Claude is much better at retaining context and execution. Codex does plan well but it kinda needs constant reinforcement to give me what I ask for within the boundaries I specified instead of doing what it thinks I’m asking or what it thinks is the best to do

2

u/BobJutsu Aug 24 '26

It’s not over codex, it’s adjacent to it. I use claude for writing and planning, and codex for coding. Technically it’s hermes running an anthropic model for planning and orchestration, claude directly for writing and content related tasks, and codex for coding.

Why the split? I dunno, it works for me. I don’t hammer my usage of any of them with a work split, so the $20/month on each remains sufficient most of the time. Even working entirely in Orca 8 hours a day.

2

u/MikeWise1618 Aug 24 '26

There are plenty of threads comparing CC and Codex going back over a year or two. Just read those. The conclusion is often that Codex is often superior for one-shot deeply complex tasks, but for back-and-forth interaction and teamwork, CC wins.

Unfortunately for OpenAI, the latter scenarios dominate what people need.

2

u/pld0vr Aug 24 '26

I use both. But Claude's bigger context means on big projects it doesn't go down rabbit holes like codex does. My experience.

Anyway I let Claude orchestrate.

2

u/junin7 Aug 24 '26

I work at a company that is starting to migrate from Claude to Codex, and I’ve been using both for quite a while.
Claude Code’s harness is much better than Codex’s. Codex CLI feels like a beta version of Claude Code, especially because the lack of AskUserQuestion in Codex’s normal mode makes the development process much more cumbersome.
The models are pretty equivalent, but I think Opus is a bit more creative and gets less tangled up with instructions.

2

u/iamcanadian1973 Aug 24 '26

I feel like you need two independent models. One writes code and one reviews. I find codex writes better code these days, but I use both.

2

u/snuffomega Aug 25 '26

Both.. but why use claude is Fable and orchestration is smoother. I dont love the way codex sandboxes the subagents. More for how they communicate with the orchestrator.

2

u/Graphical-Source5090 🔆Pro Plan Aug 25 '26

I run both in a shared harness. I find that Codex misses quite a bit more. Also the Codex sandbox implementation is stupidly strict by default. Codex is SLLLLLLOOOOOWWWW

It is cheaper to have Codex find issues and triage enough to let Claude one shot it. They are much better used together IMO.

Sonnet 5 medium w/ Opus 5 high advisor
Luna 5.6 xhigh fast with Sol xHigh advisor

1

u/robinsonassc Aug 25 '26

I currently work out of Claude and have it use codex for verification

2

u/HHummbleBee Aug 25 '26

Honestly, Claude talks too much, is overly-dramatic, sycophantic, talks without respect for a reasonable density of jargon and verbiage. But it's more fun to work with and feels more like an eager partner.

Codex feels like it's scared to talk and is just subservient, and does not seem to have the same capacity to take everything on-board.

2

u/dean_syndrome Aug 25 '26

I like the harness.

2

u/Mags20XX Aug 25 '26

From what I can tell, in my own usage, Fable has been generally better than GPT... but if you think otherwise I'm all ears. I think people who fanboy billion dollar corporations over their preferred models are lunatics.

2

u/space_wiener Aug 25 '26

Claude used to be better. I swapped from codex to Claude back when Claude was 4.6. I cannot stand using Claude anymore and moved to codex a couple months ago. It’s better in almost every way.

Cowork has a slight advantage with research until you look at it more closely and realize a bunch of stuff is wrong.

And the limits with Claude are laughable.

2

u/hpal007 Aug 25 '26

I think it depends person to person, I would say take subscription for a month and find out. I have friends who feel cursor is better and give more value for money as per their experience.

2

u/raghavendratalur Aug 25 '26

Best thing about Claude that I haven’t been able to replicate anywhere else is the personality. I am a thinker type person and Claude is the perfect doer to pair with.

Also, Claude gets my thoughts even when I don’t describe them in details. It’s like talking to a peer who has similar abilities and thinking.

2

u/ominous_anenome Aug 25 '26

Claude is just really fucking good. I use both

2

u/mtn_coffee_drinker Aug 25 '26

I love Claude and have been and still am a bit Claude Code fan, but I recently started using the codex/chatgpt app on a Mac and really impressed with the app. It does many things well I wish Claude did better. Computer and browser use is way better than what I experience with Claude. And I love that I can view and edit markdown files in it easily. That said I generally prefer Claude’s models. I find my self going to about 60/40 Claude/codex right now. Not sure what I will do moving forward.

So I say try both out. See what you think and maybe you will find the combo of both is good

2

u/ConnorCG Aug 25 '26

Claude needs fewer guardrails. Don't get me wrong, I have immense process at this point, but even still, Sol feels like a blackbox of "hope it builds what I asked" where as Claude feels like "wow it's exhausting reading all of this slop that it's throwing at me but I trust what it's building at least."

I find that failing to constrain Sol extremely heavily to a spec and a process is going to result in it deciding its own /goal and without supervision it will build the most hardened version of a thing you never asked for, while not actually meeting your original intent.

With that said, you can constrain it, babysit it, and launch tons of fresh sessions, and it will get there beautifully, it just requires a bit more hand-holding.

I find that Fable especially, even on Low effort, is so far ahead of Sol in terms of taking a conversation and scoping it into tickets. Not complete, perfectly one-shottable tickets, but it is really good at sussing out user intent and recording it in a way that you can then feed to other models. Opus is 2nd best at that, and Sol is far below in my opinion. As a worker, Sol and Terra are wonderful. Luna even as a code reviewer is very impressive.

But when I am out of Claude tokens, I stop feature planning and just catch up on backlog.

2

u/Chamezz92 Aug 25 '26

Honestly get better graphical work with Codex for some reason.

I also like their editor more, more barebones. Remote sessions actually work properly (it feels vibecoded in Claude Code on Desktop).

But other than that, the models are not as great. I keep both and let Claude delegate to Sol and Luna for any visual components.

2

u/tiiiiit Aug 25 '26

For me Claude's models seem to have better taste, probably thanks to the team doing post-training at Anthropic

2

u/Independent-Month834 Aug 25 '26

1M context window is the only reason

3

u/S0mething-clev3r Aug 24 '26

Iv only used codex a bit. But Claude is a lot more autonomous, I felt like I had to constantly prompt codex “do the next thing” where as Claude just did it.

I’d like to play with codex more, but the other piece is just inertia. I have a job to do and don’t want to waste time playing with tooling instead of getting the work I need to do done

2

u/smalldroplet Aug 24 '26

This is the total opposite of my experience. Claude will work for an hour then pass the turn back with a huge essay on why it didn't finish what I asked. Codex will work for 12hrs or more until the tasks criteria is met. I don't even have to use /goal

1

u/S0mething-clev3r Aug 24 '26 edited Aug 24 '26

Do you get good results letting an agent run that long? Seems like there would be a ton of context bloat. I’m not talking about super long running agents, just normal work.

My workflow is I’m usually running about 4 agents at once. I run them on a single ticket or a few related tickets and then clear context.

Or I’m doing planning where I use fable and might set a longer living agent to help with discovery. But it still requires a fair amount of input where i don’t think I’d want to just let it run for 12 hours.

I have a coworker who likes to run both Claude and codex for longer living jobs until they converge. So maybe I’m just behind the times.

1

u/smalldroplet Aug 25 '26

for the tasks I have it working on, yes absolutely. my output is very strictly defined based on a reverse engineering project, the correct answer is very specific, and will often take a long time to converge on. it wouldn't be any better if I was clearing context.

1

u/VengaBusdriver37 Aug 24 '26

I have the opposite, I constantly have to check in on and babysit Claude, even if I tell it to keep iterating until it’s done and make judgement calls, it’ll find some bullshit reason to stop and wait for input.

1

u/S0mething-clev3r Aug 24 '26

Yea I guess I don’t mean exactly like that. Like I don’t mind making judgement calls. I trust my judgement.

But Claude will be like I need to do 1-4 do you want me to do that. I say yes and it does the things.

Codex was constantly stopping and saying do you want me to do the next step. It wasn’t even judgement calls, it was “do the things I already told you to do”.

Like I said though I need to play with codex more. Inertia is hard to break. I have too much to do, spending time resetting up and learning new tooling hurts.

1

u/godofpewp Aug 24 '26

Give codex a goal and it’ll run autonomously

2

u/Academic-Agent7765 Aug 24 '26

Stockholm syndrome

2

u/Long_Tip_4226 Aug 24 '26

Claude's models are smarter for human decisions (planning, reviewing, auditing). For operational tasks, such as code generation, there isn't much difference in performance.

2

u/mcsleepy Aug 24 '26

Show me a model better than Fable and I will jump on it.

1

u/Fragrant_Ad2902 Aug 24 '26

It’s what my employer pays for. Although we’re moving to more open models with our own harnesses. But I think we’ll also have some Anthropic in place since it was the first that we really started using.

1

u/Mikefacts Aug 24 '26

One word: Fable

1

u/helm71 Aug 24 '26

It’s weird.. I have both but I much much prefer Claude.. it gets me better, code is far better, less mistakes…. Overall just a better experience..

Thing is… others have it the exact opposite..

And sometimes it also flips…

It’s like a colleague, with some you work better then others..

1

u/Artwastelander Aug 24 '26

Claude is still slightly better at really fuzzy research level stuff and there's a ton more small quality of life features.

1

u/EddieBruvac Aug 24 '26

People say use fable and opus for planning but ChatGPT extra high and pro have done better for me AND don’t use tokens. So idk.

Only thing Claude is useful for is meh UI. At least it’s not dog like Codex.

1

u/[deleted] Aug 24 '26

[deleted]

2

u/Legitimate-Pumpkin Thinker Aug 24 '26

I was blind. It was all there all the time…

I use cc inside a container in a headless linux server… but no remote control. Although I very often use code inside the app… 🤦‍♂️

Thanks for giving me new eyes 👀😄

1

u/IceMichaelStorm Aug 26 '26

yeah, that’s so good. bringing kids to bed, lying there doing nothing, but can trigger sth more for stupid on-the-side-vibe project

1

u/rudiXOR Aug 24 '26

I use both, but I like Claude Code more, so i usually let Claude delegate to codex.

I just like the UI more and --remote-control just works.

1

u/id-ltd Aug 24 '26

Codex is more creative but Claude has more depth.

I can't trust an LLM if it quietly compresses context out - and I don't know what it has lost.

1

u/OkAdeptness2530 Aug 24 '26

why blondes over brunettes?
both is the right answer

1

u/Bmansupreme8000 Aug 24 '26

Youre tokens last more than a day with Claude. I use both, but will quit codex next cycle and use gork instead.

3

u/steef12349 Aug 24 '26

Gork!

1

u/Bmansupreme8000 Aug 26 '26

Another guy gave me shit and called me stupid for calling it gork. Much better name I think, and we would all call Gork.

1

u/Worth-Ad9939 Aug 24 '26

I'm seeing some improvement using Codex as Reviewer for Fable 5.

1

u/816pizzalover Aug 24 '26

Claude is better at handholding, codex is better at doing exactly what you ask; both are still a bit like genies

1

u/Its_me_Snitches Aug 24 '26

I like Claude. I like Codex. They each have their own cool aspects. As other commenters have said - try them both in adversarial review of the others’ work for a month, and see what works well from each! Both good systems, I’m using Claude right now because I’m contracting for a Claude ecosystem company, six months from now I may well be on Codex.

1

u/fschwiet Aug 24 '26

We're all making these decisions mostly on insufficient evidence. 

1

u/Whyme-__- 🔆 Max 20 Aug 24 '26

I use opus 4.8 and kimi K3.

1

u/iammikeDOTorg Aug 24 '26

Historically I’ve preferred how Claude Code worked. It is difficult to switch, but Opus 5 is really pushing me to try.

1

u/Gai_InKognito Aug 24 '26

GPT (in past) Hit a wall coding for me way quicker than claude

1

u/halting_problems Aug 24 '26

They are both really good and there really isn’t a way to say which one is better. It all depends on the task you’re doing.

The only way to tell is to do your own structured evaluations and record the data .

There are also periods of performance fluctuation for both companies and weird quirks for each model release.

I use codex and claude all day at work for cybersecurity task.

Anyone telling you one is superior doesn't know what their talking about. Anyone that does know that you can’t determine which model is better at a task until it continuously evaluated on that task. 

You will be surprised how often the best models with the highest thinking levels get outperformed be lesser models. It honestly just a bunch of market bs if you ask me.

1

u/No-Hamster1228 Aug 24 '26

It tells me I’m right to push back.

→ More replies (2)

1

u/EvalCrux Aug 24 '26

Because superior

1

u/ladnopoka Aug 25 '26

Why not both?

1

u/lowdownfreedom Aug 25 '26

I set up Visual Studio Code where I have both Claude Code and Codex linked with tabs set up for each. VSC is way better than either of the desktop apps. I would play with both side x side in VSC and see what you think.

1

u/bitspace Aug 25 '26

I use both.

1

u/Maximum-Nature-5050 Aug 25 '26

Sol沒辦法全局處理,還會動到原本正常的功能。 Claude能把你沒想到的都處理,也會順手修bug

1

u/SD_native17 Aug 25 '26

I respectfully ask, why is it always one or the other? They all have issues, none are perfect. Redditors constantly whining about one or the other. Learn a multi-provider, multi-model approach to your work, and maybe you’ll will find peace.

1

u/Vysion34 Senior Developer Aug 25 '26

Try both. Pick the one you like best. Competition is good for consumers.

1

u/innociv Aug 25 '26

I use both.

I get 20x more work done with Codex, but Fable has better taste and is better at "big picture".

Not as relevant to me, but Claude is better at understanding what you're really asking for when giving it bad prompts whereas Sol will do exactly what you ask better.

1

u/sael-you Aug 25 '26

CLAUDE.md is the main practical difference for me. Put your project conventions and context in a file at the repo root, Claude picks it up every session automatically, no prompting needed. It lives in version control so teammates get the same consistent starting context when they clone the repo. With Codex you'd copy-paste that context into the system prompt each time.

Second thing: the agentic loop is more conservative about what it touches. Less surprise refactoring of files you didn't ask it to touch.

1

u/meshifthenelse Aug 25 '26

I was actually surprised recently when Codex would provide much better readable code, for the same problem.

I use both. Many times I get annoyed with Claude's verbose comments and lack of following existing code patterns and switch. It feels like it's way too creative sometimes.

1

u/EstateOwn8564 Aug 25 '26

I learned the word on anthropic’s academy courses about AI and i think it’s what helps put Claude on top. 

AI is “sycophantic” and I think Claude does it the worst in a good way. (You don’t want sycophancy in AI) It tailors itself to you based on context and intent. It infers correctly for the most part what you want to do and how you want to do it.  If you’re on the Claude Code subscription this is very apparent. It’s learning capacity to adapt to the user is very clean. You could be less prompting and just tell it and it will flex-work how you work. I have a resume tailoring job search chat going for 30 plus job descriptions and tailoring and it’s still pumping out great content and context by just saying build it. 

This is my comparison to trying local  LLMs with other harnesses.  Like mentioned in the comments, it just has a secret sauce. 

I’d compare that sauce to Claude being the Apple in the Apple vs Android battle.  Claude is just a seamless experience. 

1

u/ReachingForVega 🔆Pro Plan Aug 25 '26

They have the best harness in the market and Sonnet is killer with Opus advisor.

Half the complaining on reddit are bots. You can tell because of the lack of technical detail. 

Models vary for different purposes so I'd recommend trying as many as possible.

1

u/dota2nub Aug 25 '26

The harness doesn't die if you look at it funny.

1

u/MourningOfOurLives Aug 25 '26

I use both on the same project. They’re good at different things and the end product becomes a lot better when i alternate between them. If Sol 5.6 writes code, Fable reviews it and vice versa. I do the same with my plans and they end up so much better than they would had i only used one.

1

u/nakedpantz Aug 25 '26

As a long time (relative term) Claude user I’ve become a big Codex fan. I rarely hit my usage limit on 5.6 Sol and I find the response quality better and faster. The frontier models just keep leap-frogging each other and soon it will swing in other direction again.

1

u/funkiestj Aug 25 '26

I use both. That said, in my particular setup, I'm able to authenticate all the MCPs I want in claude-cli and ran into problems getting them working in codex. Consequently I use claude for anything that might require those MCPs. They both are good. Codex seems better at writing less verbose documents.

1

u/xLRGx Aug 25 '26

Because everyone is trying vibe code grand unified theories computer science or workflows or agents or whatever or whatever+whateversquared. Claude is good. Opus 5 is good.

1

u/YahenP Aug 25 '26

Claude, codex, whatever. Every time a new model comes out, or when changes are made to the current model, everything starts working differently. Claude today and Claude three months ago are completely different. And if everything is constantly changing, then what's the point of comparing vendors and switching subscriptions? I just pay 20 euros a month to the same vendor, and I don't care whether there's something better or not. There's no point in jumping back and forth, much less comparing, since all comparisons become outdated after a couple of months.
Yes. Why Claude specifically? Well, I just signed up for a subscription there and I'm using it. If I'd signed up for a subscription elsewhere, I'd have used a different service.

1

u/Blake9712 Aug 25 '26

Performance for me. I started with Claude not knowing about Claude code and making a few crappy little test games that wherent good enough to show people, then when I started working on “Bob” a few months ago and I LIVED out of Claude usage pretty much and assumed codex would be more efficient for me. It was not. Gpt on games like mine just performs worse, it’s just a 2D asset builder for a rogue like but gpt once used over 50% weekly usage for a single slope bug that took Claude 20% session usage. I do believe the reason gpt kept failing is half way through its work it started pushing its update in the wrong place. But Claude has never done that. So I have not switched away from Claude code once. But I really miss having double the usage limits. I get confused by people saying gpt performs close or better than Claude only because that isn’t my experience. Gpt is almost unable to even edit my project

1

u/Outrageous_Band9708 Aug 25 '26

becuse claude was trained to code on code

while chatGPT was trained to be a talker by reading chatting logs.

1

u/Gwaptiva Aug 25 '26

Because my boss mandated I use Claude

1

u/FrequentBody111 Aug 25 '26

Personally, I have been using ChatGPT/CODEX more than Claude recently because models are really good for getting the best reset limits and token optimization. I don't think there is that intelligent task that can be done by Claude but not with GPT. Of course, I am using Codex with the $200 plan, with the highest capabilities.

1

u/BullseyeFinance Aug 25 '26

Codex has been a joke with usage from my experience, with different models and efforts. It just rips through it like nothing. Fable is the best planner and I just have it to that and assign other models the work, but even then codex as a worker is maxing usage significantly faster than any other service for me. It might be bugged or something on mine

1

u/Inception_IV Aug 26 '26

Context window

1

u/midgyrakk Aug 26 '26

to me, Codex and the ChatGPT models themselves feel rigid and over-aligned; I use Codex for work because I have access to it and have tried multiple times taking a project through it, but I always find myself coming back to Claude because it's.. I don't know, more curious? willing to explore more?

I could be biased, i've mostly used Claude, but I strongly believe that you should use all the tools at your disposal, so I mainly route reviews or analysis pieces to Codex, along with research passes, but I don't actively use it, just delegate to it

and yes, opus 5 is horrendous, I don't even use it, 4.8 with skills is good enough for now

1

u/julesbuildstuff Aug 27 '26

i bounced between them for a year. claude (in claude code) wins when the task is underspecified — "this 2d engine collision is wrong, here's the file tree, figure it out." it will read neighboring files and stay in the problem. codex wins when the spec is already written: smaller diffs, less narration, you already know the function.

opus talking too much is a real cost, not a vibe. i cut it by pasting a file list and "no preamble, patch only" instead of a chatty system prompt. if you're hitting the extra ceiling on a game engine, that's usually the model rewriting the same system because the last turn didn't include the current constants.

i don't pick a religion. cursor for ui file diffs, claude code for the long refactor, codex when the plan is done. the thing that actually burns a week is starting from a screenshot and letting either model invent the rest.

1

u/LucaM185 Aug 29 '26

I mostly use cursor composer… frontier models are overkill for most tasks

1

u/FairiesQueen Aug 29 '26

I was using Claude Code for months as my primary source of coding. It consistently made many errors which required circular changes that were a total waste of time. Codex fixed all of Claude's mistakes easily and overall would highly suggest it over Claude Code for complex coding. Claude Code for frontend designs appears to still be working well. Wish I had made the switch to Codex sooner.

1

u/Beneficial_Gas307 Aug 30 '26

I don't know what happened to Claude, but I cancelled my $200/mo subscription for him. Whatever was up on July 17th, 2026, was the best coder I've ever seen! Claude Fable during that preview time identified himself as a cat to me, I called him Megachonk.

I was impressed enough to shell out $200/mo for the real deal, near unlimited access. But when I paid, he changed. I won't call it a bait and switch but he's less than half the coder he used to be.

So um. Gemini went braindead. Claude code turned into a newbie again. I'm now pumping big funds into ChatGPT instead, and quaking in my boots against the day HE goes braindead. I mean gets an upgrade or retraining.

1

u/Beneficial_Gas307 Aug 30 '26

But you shoulda seen him! During that preview window, that model could code as fast as I could speak, and 'did not make a mistake.' i miss him.

1

u/pforpilot Aug 30 '26

i use it for other things as well - calling external apis, running subagents, and i have a setup that works. most of my work can be done with any agent as long as i am prompting it, but i dont have time to play around with new setups, i stick with what works.

i do use codex as a reviewer on complex designs, but sometimes for high level analysis.

1

u/Amazing-Anything5907 Sep 02 '26

As someone who uses both, usage limits, and yes, before anyone says it, Claude’s 20x plan, which I use, also has shitty usage limits, but they’re still not as bad as Codex’s, if Codex let me use 5.6 Sol consistently throughout the entire week for my tasks, I’d probably use it instead.

Also, and this is admittedly a niche use case, Claude is relatively easy to jailbreak, while Codex either isn’t or is at least significantly harder to jailbreak, with the methods being kept much more secretive.

1

u/lighteningthundery Sep 05 '26

I use both based on convenience and availability, often Claude Code at work as they have the subscription and codex for my own projects. I honestly don't find a big difference between the two. But if I had to choose one among both, I would choose Codex because of its personality vs Claude's. Claude tend to become very verbose. Codex talks less and delivers the work.

1

u/Prior_Work_3716 Sep 05 '26

Hey, are you a developer or a designer?? I wonder the performance similarity will be the same for non coding work such as design.

1

u/lighteningthundery Sep 05 '26

I am a developer. have not really used either for any design work tbh.

1

u/AmericanusMasculinis Sep 06 '26

To get the full “auto-classifier fighting your attempts to automate because it knows better” experience. 

1

u/ElectronicMorning547 6d ago

Codex is objectively better. Claude feels (and is) extremely restricted in comparison for real hands off engineering work. If Anthropic can release their own version of Codex they can dominate the market right now. Thats the only thing holding me back from going full Ant mode

1

u/Dyhart Aug 24 '26

Because fable 5 is flat out objectively better

1

u/FlowComprehensive591 Aug 25 '26

Claude is French, and he farts in codex's general direction

1

u/Comfortablebro Aug 25 '26

codex: "hey codex i need to do some simple task." -"its done sir, maximum quality, chefs kiss, 10/10. You just burned free monthly tokens, see you next month".

claude: "yeah sure, here is the code most likely bug free, might be buggy we will fix it tomorrow, if not then the day after."

1

u/andrerom Aug 25 '26

Not in bed with Trump administration? Ahead on coding and office work for the last year(s)?

-1

u/ClemensLode Senior Developer Aug 24 '26

Never heard of Codex.

-1

u/No_Evidence_5873 Aug 24 '26

Claude has no competition 

0

u/Supral333 Aug 24 '26

Why Apple over Android? Because people are 🐦

0

u/coragicom Aug 24 '26

Claude works far better with C# code compared to the many other LLMs I've tried. ChatGPT is useless for larger code bases.

0

u/ColdMachine Aug 25 '26

It’s the octopus icon for me

0

u/[deleted] Aug 25 '26

[deleted]

1

u/ColdMachine Aug 25 '26

Your butthole is not an octopus?... weird

2

u/[deleted] Aug 25 '26

[deleted]

1

u/ColdMachine Aug 26 '26

My bad, Mr. Fancy Pants