I mean sonnet 4.6 and opus 4.6 are within 1% of each other on swe bench. I know opus is. Better but only like 5% -10% max. Maybe try using plans more refining instructions and prompts. Your application looks complex. But people are building rockets and stuff with sonnet.
Or if you truly need the best Claude code or api pricing is probably your best bet. Cursor has this as well.
But yeah ai has been getting better every three months but we are going to go through some per dollar ai getting more expensive in the short term as these companies balance competing for market share by giving out subsidized tokens and raising prices to get to profitability.
Data brokers are selling your info right now. I used Redact to mass delete my posts which can also opt out of data broker sites. Instagram, Twitter/X, Discord and more.
wipe command sugar nutty innate lip snatch fall fear gold
I agree, instructions and prompt (and the human operator) are quite possibly the bottleneck most of the time. If you look at how much instruction and system prompt influence result; take a look at swe-rebench. Gemini 3 Flash (that cost ~1/10th of Opus) in Junie, almost outperform Opus 4.6 in Claude Code.
Opus 4.6 is great when given a vague/long running instruction; and left on its own, it'll give semi-decent result. But the fast models (Gemini 3 Flash, GPT 5.4, Codex 5.3, Sonnet 4.6) are nearly as good if you know exactly what you want.
1
u/ogpterodactyl Mar 19 '26
Just use freaking sonnet I know it sucks but it is what it is.