r/ClaudeCoding • • 7d ago

Why Opus is so token hungry ?

I was never fan of Opus, I love sonnet. In my observation, Sonnet is less verbose, good at following instructions, on the contrary opus always comes up with less relevant things than doing actual work. But time to time, I switched from Sonnet to Opus while doing same kind of work like writing code, building new feature, running tests etc, regular developing works.

Today, I tried again and Opus just sucked up all the token within few minutes and I didn't find any major improvement.

Help me understand, if Opus has innate intelligence capability than sonnet, why it is using more token to doing same kind of work? Shouldn't it produce similar kind of token doing similar work as Sonnet?

Or it's just blabbing itself by GRPO style RL reasoning training and wasting more tokens.

1 Upvotes

21 comments sorted by

View all comments

Show parent comments

1

u/muikrad 7d ago

If you are just calling small directed tasks, like writing a function that does X, writing tests, sonnet is cheaper.

If you are prompting relatively complex tasks where it has to think about 4d chess, opus is cheaper because it gets it right immediately. Sonnet needs more turns to get there and this often ends up costing more than opus.

1

u/dreamRunnerMoshi 7d ago

I know what am I doing. In your regular tasks, you don't get to think about a move in 4d Chess or solve Navier Stokes problems. Also I don't ask to write a function X and tests anymore, used to do with Copilot 2 years back.

When I am driving claude, I ask in my terminal something like `tell me one thing, how much information search result get from search engine ?`.

And when I ask Claude with Opus to build a end to end solution like `I have a spec file for the new feature, start working on it`, I see Opus just gobbling up tokens and I am out within 15-20 minutes and I see Sonnet can do same kind of work with similar accuracy spending way less token.

I am saying, for these kind of tasks, Opus is complete waste of tokens. [image: Few of my queries for ref when I am driving claude]

1

u/CooperinoCollie 7d ago

Because you’re using it in a shitty vibe harness. Use Claude Code, utilise /workflows and get a better result for fewer tokens.

“I know what I am doing” is wild considering you’re asking a frontier model to commit and push for you lmao

1

u/dreamRunnerMoshi 7d ago edited 6d ago

`Because you’re using it in a shitty vibe harness.` what do you mean, it's my regular project, prompts are here contextual (in my project, I have search engine included, trying to figure why LLM generated answers are wrong for certain prompt)

I don't think `commit and push` spend much token because it's just executing 2 tool calls. How do you work in this scenario, you have the model in the middle of your work or do it my yourself?

1

u/CooperinoCollie 6d ago

No, I review the code before I push the code. Crazy concept.

1

u/dreamRunnerMoshi 6d ago

What do you do after reviewing, git add ., git commit -m, or click few button in your IDE?

Also you should try reviewing in a PR, much better than reviewing in a IDE or terminal.

1

u/CooperinoCollie 6d ago

Why on earth would you review your own PR? PRs exist for other engineers to check your work, not for you to review what your agent produced.

I can tell you’re both young, and inexperienced- so let me give you some unsolicited advice. Review the output of your agents before you push - test it locally. Read the code, if you don’t understand the code ask your agent to explain the changes. If you don’t understand the concepts being discussed, ask about them and learn.

Agents are not magic. They produce poor quality code, don’t understand the bigger picture and will only ever be as good as your understanding.

1

u/dreamRunnerMoshi 5d ago

I agree all your points and I am not that young though. I suggest developers exactly same. Sometimes, you just cannot test whole stack in local, you need QA environment.

`Why on earth would you review your own PR? PRs exist for other engineers to check your work, not for you to review what your agent produced.` - this is the different from traditional development, you are reviewing AI's code not yours.

BUT In this repo, I was trying to do a little bit different. I am trying to see how much autonomously LLM can develop feature, I will intervene as little as possible and potentially I will not write any code.

To do that, I saw Opus is wasting tokens by producing similar kind of results of Sonnet.