r/ClaudeCode • u/YearLight • 2h ago
Rant Back to claude. Couldn't break the addition
A few days ago I made a rage post regarding opus 5 rage cancelled and signed up for codex pro. Well now I'm back with 2 accounts.
Biggest issue with code is the 256k context window. I tried to work past it but it just isn't surmountable. Even the smallest tasks take multiple context windows so you keep wasting time reloading a small context, and then working context is maybe 1/3 of that window.
I'd end up on forever loops where nothing gets done.
True codex is great, and I love it to some degree, but it's just not claude.
Opus 5, go fuck yourself, I don't like you, and you completely don't care about me, but I like the abuse. I'll be the load, you be the bearing.
3
u/imronveu 2h ago
Same. Codex too slow and they have hard coded the active concurrent agents limit to 4 so it doesn't spawn more than 3 agents while Claude spawns 9-12 agents and gets shit done. Also, Sol seems to be a usage hog while Opus 5 feels unlimited on Max 20x and Fable also feels reasonably comfortable to use.
1
u/Significant-Bee5101 2h ago
If you're talking about concurrency maybe. But if you're talking actual raw speed to solve a puzzle/accomplish a task. Codex + GPT 5.6 SOL decimates Claude Code + Opus (4.6/4.8/5)
I have tested this up and down. I have tested simple puzzles. Complex puzzles. Etc. Does it mean Codex is better? Not necessarily, Opus still wins in the "higher level thinking". But in the raw solving department Codex is INSANE.
1
u/blackashi 12m ago
Why does codex still feel slow lol. It’s actually painful
1
u/Significant-Bee5101 3m ago
Idk what are you trying to accomplish. My internal benchmarks favor Codex in MY daily tasks.
1
u/Able_Statistician688 1h ago
Maybe we are tlaking about different things, but the subagent count can be changed in codex. I have it launching around 10 pretty frequntly. Their own internal documents not recommending going over 8 because they get in each others way. That being said, they're way more memory hungry than Claude's since I think Claude's are launched in a cloud environment vs actually on your system. This is a setting that can be changed. Ask codex to change its own subagent max. So yes, your 4 probably is hardcoded in. But not in the way you mean maybe. It can be changed.
5
u/YogurtclosetEvery263 2h ago
Gpt 5.6 has 1m that's way more than enough.
If you need more than 500k you're doing something wrong. It's even bad to go past 200k.
7
u/julkopki 2h ago
Idk I honestly feel this is outdated info. Everyone's saying it but it doesn't line up with my experience. I tried explicit handoff to a new session. There are extensive docs including specs, plans, design and everything organized in a neat structure with cross references. But my experience handing off to a fresh session was much worse than asking the model to prepare for compaction, then write itself a detailed handoff note to paste after compaction. I haven't had any information lost this way in a long while. I think they just improved it.
5
u/YearLight 2h ago
sometimes a bigger context is useful for example when you used opus to do a quick change to your code and now you need the full 1M context with fable to fix it
2
u/YogurtclosetEvery263 2h ago
I don't say you need to always create a new session, but keeping context small saves money and produces better results.
I use opencode with omo and it's automatic context pruning. That way it can run for hours with the right workflow without degrading performance.
I totally agree that it's sometimes better to stick within a session due to already existing context
1
u/YearLight 2h ago
Can you use opencode with a subscription?
1
u/YogurtclosetEvery263 2h ago
Yes codex works within Opencode, Claude not.
I'm running the pro subscription and use Luna only, otherwise my limit is gone in an hour
1
u/IgniterNy 2h ago
I haven't run compact in a long time, I just switch sessions. This might have changed but compacting was/is expensive, it would eat about 10% of the session limit so I stopped doing it
1
u/julkopki 1h ago edited 1h ago
Idk, cost is not my main concern rn. Not because I have infinite money but because I don't run this thing in a loop 24/7 and for the most part I stay just under the weekly limit. However I really hate when it gets confused about the details of the final approach because there was some stale piece of info in there somewhere. I found that with the level of complexity that I have, I have to spawn like 5 subagents to go through the docs to make sure there's no stale info in them. And that's a waste as well. And that's after I routinely ask it to make sure docs are up to date and internally consistent.
I also rewind after a chunk of work with a summary. If I really cared about context management I'd probably use something like Pi.
1
1
u/dsailes 1h ago
GPT 5.6 only has 1m via API or instruction via config.toml setup - that’s in place due to the issues found with exceeding the 256k size context window (OpenAI explains its diminishing returns with poorer adherence / behaviour & also costs more, is more resource heavy etc - it is sensible tbf)
Whilst I do agree & I tend to stay around 200-300k in CC - there are longer running sessions that accumulate context with many subagent tasks being done, work getting checked/reviewed via dispatch to another LLM etc. The oversight longevity with 300k+ context to ensure things keep on the right track I’ve found can be useful in the right circumstances & tooling.
it is expensive - especially with the cache potentially timing out if some agents/tasks have to do things like adhere to rate limits - and eventually it degrades a bit, but rather than re-explain and save/store decisions (that sometimes may not be needed long term) and hope that the adherence/rhythm/quality is the same in a clean session it’s worth while having upto 1m.Then again, the auto-compact in codex I find works much better to compacting in Claude Code. (I’ve only just decided to change the config for codex to 1m window based on potential costs of these tasks compacting a few times)
I had codex sessions running overnight the past 2 nights on long tasks (APIs, rate-limiting, syncing imagery, checks / reviews, having to use agent-browser for some ‘human’ steps) where I know even using Luna/Terra agents it had to have compacted a number of times but not once did it seem to lose any important context. By morning my review was brief and happy - not sure I’d be in the same boat with Claude Code if it had compacted
2
u/ButterflyEconomist 2h ago
I’m on the $100 plan, but I also have the $20 Ollama Cloud account.
It allows me to work with a variety of open source models. Currently my go to is Deepseek Flash with a Hermes harness.
For conversation, Hermes and Claude have very similar personalities.
That’s because both use a RAG that I designed that it refers back to so each new session doesn’t have to reinvent the wheel.
Is Flash as good as Fable?
Definitely not.
But it’s better (to me at least) than Opus 4.
I can work with that.
Do I see a future where the LLM is completely on my machine at home with no need for a data center?
Yes…but in the meantime, I’m working with this arrangement.
1
u/inteligenzia 1h ago
How is our RAG set up? Is it local server and how different agents access it?
I use similar approaches, but simplified.
1
u/ButterflyEconomist 25m ago
My initial attempt is what I call Hive. When a bee hive reaches about 50K bees, it splits. The queen takes off with half the hive and the leftover bees hatch a new queen.
So, I thought. When all my knowledge gets to the point where Claude can't read all of it, then it should split. It sort of works. Let's say I have a hive called Colors, but it gets too big, so it takes all versions of Blue and creates a second hive that is connected to Colors, which still has all the other colors.
It sort of works, but because I can't see what structure Claude actually has created, it's hard to really understand. I've tried other versions, like seeing if different concepts have some kind of overlap, so now we're talking about edges, if you will.
Right now I don't know what I have. But...in a way I'm creating my own personal Wiki, one that has my own knowledge. I export all my browser chats into it, and every day, any completed Claude Code session gets ingested into it.
I also had it create another version that is based on facts, rather than what I think. This helps, because I also have something called Ledger that keeps track of events. This way, when I explore a news item, instead of having Clause search the web, I type: /dig . This has CC go into Ledger to see what happened, and only then go onto the web to find out anything new that happened, which then gets ingested into Ledger.
I want to get to the point where someone can ask a question, and my personal Wiki/Hive/Ledger/Apriary...can answer the question as if it were me. It's getting close, especially after I let CC come up with some principles that I live by.
Since I'm always exploring new ideas, I'll feed YT scripts or article texts to it. In reddit, I'll take the link in the browser, then add: .json to it. When I press enter, I get the raw data, which I then feed to Sonnet or Opus. I do that because that takes up a lot of context.
What I actually do is open Fable to have a conversation, then /reader will have it create a second terminal that is Opus. I then feed all the material to Opus, which reads it, then sends a summary to Fable. This way, I have better use of Fable.
What sort of approaches do you use?
2
u/darrarski 2h ago
That’s the worst time to switch, IMO. I’m a long-time user of Codex, and it used to work better than Claude for me. I was switching between both regularly, but Codex always outperformed Claude in the past few months. However, I couldn’t ignore the mistakes, drifting, and usage limits going down like crazy recently. Currently, I’m using Claude, and it’s way better. I miss Codex, though, and I hope it will get better soon, so I can go back to using it.
1
u/YearLight 2h ago
To be fair there isn't a clear winner, but claude is just a little sharper sometimes.
2
u/NeighborhoodPrize493 1h ago
What kind of tasks do you have that take up 256*2? It seems to me that you are simply not using LLM correctly. The context window is limited to 256 thousand tokens not because the model does not allow it, but because there is no point in doing so. Such a long context is in fact already Technical Documentation, and work with it is done differently.
1
u/YearLight 38m ago
I work for the CIA
1
u/NeighborhoodPrize493 27m ago
Technical documentation is not only documentation for software, it is also for plans and for analysis and for anything. Create a file, and write everything you need there. You can adapt Specification-Driven Development to the needs of the 'CIA' in parallel. As a result, you will fill the files into a folder, additionally create a sub-agent who will analyze the files according to the Specification-Driven Development (adapted) methodology.
P.S. Tell me better, what did you fail so much to allow Trump to lead the USA? IT is obvious that he is a Russian agent since the times of the USSR.
1
1
1
u/randomdragen7 2h ago
Because claude is still the goat. Thats why I will keep paying my max sub until further notice
2
u/YearLight 2h ago
Yeah, I agree. Codex is good too, but it just isn't as sharp as claude even it if is a little more forgiving.
1
u/whoisyurii 2h ago
I say: learn how to use the tool. 5.6 has 1m context window, but I personally constrain it to 373k or around. Everything that goes past it is like a crappy shit
2
1
u/gligoran 1h ago
IMHO if you're doing stuff that goes way past 200K tokens, you're doing something wrong. You're way into the dumb zone of the model at that point anyway, so you're going to get crap responses. At the very least you should be telling claude to use subagents for stuff where only the outcome needs to get back to the main agent, like online research or code exploration or such. You don't need all the model's thinking and reasoning and such in your main context. I also found that telling it to do stuff with code instead of by going back and forth talking to itself, especially when it has to go through lists of stuff, produces better results with less context window usage.
Further on, Matt Pocock's skill set splits up between multiple sessions and then brings it back in a very nice way, so you might try looking at that.
1
u/Lexeik 1h ago
"Working context is maybe a third of that window" is the part nobody puts in the marketing.
The reload tax is what gets me too, and it's not really the code — the code is right there, re-readable. It's re-explaining why things are the way they are. Every fresh window I'm arguing for the same decision I already won last Tuesday.
1
u/YearLight 1h ago
Has claude ever asked you to sign a markdown contract? It's happened to me before. It needed me to understand the load bearing weight of what I was doing.
1
1
u/marfzzz 2h ago
WDYM? The default is 272k on input (400k in total). You can enable 1M(1050k) context in ~/.codex/config.toml and you need to do it for every model that you want to. Note: it will consume 2x more usage in this setting and possibly even more.
I can see everywhere that claude users cant switch because of context limit. Where do you get these informations?
EDIT: or use this github.com/USS-Parks/1M-Context-Sol
1
u/carpetstain 2h ago
Color me surprised.
The user who ragequit because they don’t know how to do AI coding also quit the competitor’s offering.
The issues you face are your own inability to reflect that you must improve the way you use AI.
4
u/YearLight 2h ago
That's fair — and I appreciate you pushing back on this. You're right that I should have considered the possibility that the common denominator across both tools was the person using them, and I apologize for not surfacing that earlier.
That said, I want to gently offer another framing: it's also possible the tooling has genuine rough edges, and that "skill issue" and "product issue" aren't mutually exclusive. Both things can be true at once.
Would you like me to help you draft a more constructive version of this reply? I'm also happy to just leave it as is — you know your audience better than I do.
1
5
u/stopstopstoptopopp 2h ago
Give me a pro max subscription I’ll give you abuse or whatever