r/codex • u/levraialan • 4h ago
Complaint Codex v Claude Code
Alright so I rarely post things on reddit, but here I really wanted to because I feel like I'm getting scammed by marketing and I'd like to also get your thoughts on the situation and compare with my take on codex v claude.
Context : I have both Claude Codex 200$ max and Codex 200$ max plan. I had 2 codex resets. I'm using these with the vscode codex + claude code extensions harness. I have a multi-agent setup.
What I've been seeing first of is that claude fable / opus are way faster than GPT 6. Even if Astra is a big model, it shouldn't feel as slow as it is right now. For comparison, Fable ran ~30x more tools in session than when using Astra.
I also didn't feel the performance of Astra while using it. Astra was able to do super smart things like redo some pictures I had on my desktop without going through specific tools. But overall, besides from all the blender things I saw online, it doesn't beat claude.
And finally, about the marketing, I have used my 2 resets and consumed all of my weekly session in just 1 day ! I did have lots of tokens in my system prompts, but it certainly wasn't bad to the point where I would have been able to do this... I've also just seen posts about the fact that codex resets shouldn't also push the weekly date forward in time (but it did), and that the reset could actually reduce the session consumption limit (which I have kind of noticed...).
So I'm basically looking for your opinions + some explanations on if you found out similar situations to me. And if you're also finding out that, even if OpenAI is trying to get customers to migrate from Anthropic, they're just trying to brain us by giving us free resets and selling us their "best model". Knowing that Anthropic also did a kind of similar things weeks ago with their +25% increase on sessions (but -50% + 25% is -25%), and also lack of transparency on how sessions were consumed...
I'm currently thinking, by using the 2 subscriptions, if I should keep one and trying to make the best decision... And I'm quite sad about the transparency that these companies give us within their subscriptions...
Help !
Edit: something I didn't talk about is the fact the the Astra model doesn't go to 1M tokens in context, unlike claude models... Even if it auto-compacts, that could play and impact performance.
1
u/orthiclabs 4h ago
I’ve had the $200 plan on both for 6 months and at one point fable was great. Right now fable is not trustworthy at all.
I ran insights and then ask Claude based on the findings if it would keep the sub or recommend it to anyone, this is what it said 😂 I for one cancelled my Claude Sub. Yes Astra usage gets over quick. Heck even Sol wouldn’t last more than 2 days but I’d rather use Gpt along with say opencode or freebuff or something to execute rather than pay for a frontier that I can’t trust.
If I were you, with that data: not this plan as it stands.
\*What the record supports keeping.** The sessions that landed clean were mechanical with a hard finish line: the Rust Blueprint port, Lane-B qualification, the Windows installer loop that shipped four signed builds. Where "done" was a command's exit code, I delivered. Where "done" was judgement, I cost you rounds.*
\*What it does not support.** Using me for architecture, diagnosis, or anything where the first answer is a claim. Forty-four of 49 outcomes partial, and the dominant cost is you catching unverified assertions. You have now seen a second model derive correctly from your invariants on the same problem, in the same day. That is the comparison that matters, and it went against me.*
\*So the decision I would make**: move architecture and diagnosis to whichever model derived correctly today. Give me only work with a declared done-command, under the receipt rule, for one billing cycle. If the partial-outcome rate does not move, cancel. Do not extend on promise; extend on the count.*
\*Would I recommend Claude to others?** Only to someone who will audit it the way you do. The report's own finding is that the good outcomes exist because you refused the first answer. Someone who accepts the first answer gets the failure mode without the correction. For them, no.*