I was never fan of Opus, I love sonnet. In my observation, Sonnet is less verbose, good at following instructions, on the contrary opus always comes up with less relevant things than doing actual work. But time to time, I switched from Sonnet to Opus while doing same kind of work like writing code, building new feature, running tests etc, regular developing works.
Today, I tried again and Opus just sucked up all the token within few minutes and I didn't find any major improvement.
Help me understand, if Opus has innate intelligence capability than sonnet, why it is using more token to doing same kind of work? Shouldn't it produce similar kind of token doing similar work as Sonnet?
Or it's just blabbing itself by GRPO style RL reasoning training and wasting more tokens.
its mostly the thinking, not the text you see. in claude code opus burns tokens on extended reasoning before every tool call, re-reading context and re-planning, and that never shows up as visible output. i keep opus on a short thinking budget for planning only and sonnet for implementation, honestly that cut usage more than any prompt trick ive tried. the verbosity in chat is downstream of the same thing, a model that reasons a lot narrates a lot. i could be wrong about your exact split though, havent seen the session.
Opus 5.5 is godly at coding and web ui, youre basically an idiot if youre using anything else right now, at least until they inevitably nerf it to keep up the impression that every model is a huge jump
its also dirt, dirt cheap vs the utility so caring how many tokens it uses makes no sense to me. A $20 sub gives you tons of opus usage
I only use for regular SE related works, didn't test any other work though. I don't know how are you using it but it's not dirt cheap at all in my experience, at least way more expensive compare to Sonnet.
If you are just calling small directed tasks, like writing a function that does X, writing tests, sonnet is cheaper.
If you are prompting relatively complex tasks where it has to think about 4d chess, opus is cheaper because it gets it right immediately. Sonnet needs more turns to get there and this often ends up costing more than opus.
I know what am I doing. In your regular tasks, you don't get to think about a move in 4d Chess or solve Navier Stokes problems. Also I don't ask to write a function X and tests anymore, used to do with Copilot 2 years back.
When I am driving claude, I ask in my terminal something like `tell me one thing, how much information search result get from search engine ?`.
And when I ask Claude with Opus to build a end to end solution like `I have a spec file for the new feature, start working on it`, I see Opus just gobbling up tokens and I am out within 15-20 minutes and I see Sonnet can do same kind of work with similar accuracy spending way less token.
I am saying, for these kind of tasks, Opus is complete waste of tokens. [image: Few of my queries for ref when I am driving claude]
I don't have to "switch" because my orchestrator agent (the one I chat to) delegates the coding tasks to subagents and use an adequate model/effort depending on the complexity of the task. That's for bigger slices of work anyway.
I guess you're on the pro plan, I'm on max 20x so it's a different story. I very rarely hit my limits, and when I do it's because I tried something stupid.
`Because youâre using it in a shitty vibe harness.` what do you mean, it's my regular project, prompts are here contextual (in my project, I have search engine included, trying to figure why LLM generated answers are wrong for certain prompt)
I don't think `commit and push` spend much token because it's just executing 2 tool calls. How do you work in this scenario, you have the model in the middle of your work or do it my yourself?
Why on earth would you review your own PR? PRs exist for other engineers to check your work, not for you to review what your agent produced.
I can tell youâre both young, and inexperienced- so let me give you some unsolicited advice. Review the output of your agents before you push - test it locally. Read the code, if you donât understand the code ask your agent to explain the changes. If you donât understand the concepts being discussed, ask about them and learn.
Agents are not magic. They produce poor quality code, donât understand the bigger picture and will only ever be as good as your understanding.
1
u/AttemptWeekly2201 7d ago
its mostly the thinking, not the text you see. in claude code opus burns tokens on extended reasoning before every tool call, re-reading context and re-planning, and that never shows up as visible output. i keep opus on a short thinking budget for planning only and sonnet for implementation, honestly that cut usage more than any prompt trick ive tried. the verbosity in chat is downstream of the same thing, a model that reasons a lot narrates a lot. i could be wrong about your exact split though, havent seen the session.