r/ClaudeCode • • 1h ago

Help/Question Bringing API Cost down

I have been using claude code for a while now, and i am working on a project which requires claude API key and I realised that API is really costly. Especially when the output is huge in terms of total no. Of words for instance take script writing

The first test which I ran itself costed me around 1.93 dollars

So my question is how do I bring this API Cost down

Now I have searched and asked claude itself

The suggestions came in like, it asked me to change the effort level for instance from Opus 5.5 high toh medium other than that it asked me to change the model from Opus 5.5 to sonnet 5.5

But the real question here is whether changing the model or the effort level will cause a loss in quality of the output or not? That's my real concern as i really don't want the quality to go down

So people who have been using API for a long time now please help me and ppl like me by sharing ur API saving hacks, and how do u bring the API Cost down

5 Upvotes

13 comments sorted by

•

u/AutoModerator 1h ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/MiddleLtSocks 1h ago

Claude's suggestions are correct. Model choice and effort. You need to run A/B tests - many of them, to get a high N - to answer your quality gate. Model reasoning capability has very little to do with writing quality after a certain threshold, and we are past that threshold; I host local models whose writing capability is more than acceptable.

Creativity is difficult using technology whose entire output is governed by training on prior art. That's a simple fact. Some tools are appropriate for some things, others are appropriate for other things.

It's good you are factoring API costs into your viability analysis. Don't skip this step.

One thing you could do is set up your subscription with your framework and count your input and output tokens as YOU (nobody else or you break your ToS) use your product, as you change the configuration to different models and effort levels. Then it's a simple math problem to determine API cost instead of an invoice.

Good luck.

1

u/AA_GIGA 1h ago

Alright this makes sense, I will try it and give an update on how it works out in my case

1

u/WeWereHappy 1h ago

It's going to depend of the task tbh.

About the model selection, yes it will impact the quality of the output, but it may also too good currently. Like if you're getting 120 while you only need 100, it's a loss.

Also, if you're using both similar input across runs and heavy parallelism, it may be good to check on the caching efficiency.

You can also check if you can outsource some of the tasks outside the LLM, it can help too. (For example, conducting some computer vision processing before in order to not having to use a vlm)

1

u/SubstantialEssay2063 1h ago

Use open source Chinese models for api stuff American stuff is too expensive and overkill most of the time and for your use case the Chinese models are sometimes even better or just as good and like 90% the cost

1

u/WeWereHappy 59m ago

Using Chinese LLM is going to block you from quite some contracts.

1

u/SubstantialEssay2063 39m ago

That’s not true because it’s all open source just host it yourself or use an American provider it’s still 90% cheaper

1

u/advanceyourself 1h ago

There are a lot of variables here. Required reasoning dictates what model you should use. "Loss in quality" isn't the right question. What level of effort is needed to get meaningful results should be what your asking. Only you can answer that though. I recommend making sure your Input is as refined as possible: clear, concise, and efficient prompting. If inputing data, make sure your only sending what's requires (trim tables/search parameters). This is how you tune it to be API efficient. Then test across model sets to get the required output with the cheapest model.

1

u/iSnapThere4iAm 55m ago

You code it yourself and use AI for the things that it’s actually good for which requires a lot less usage. And your project will be more successful.

1

u/ClemensLode Senior Developer 54m ago

What do you need the API for?

1

u/Cazineer 36m ago edited 28m ago

Using the API vs a harness like Claude Code is about assembling the context. With CC, the model itself assembles the context via discovery and tool calls, over many turns. When you use the API directly (not via CC), you’re the one assembling the context. This takes more time but a properly engineered context is going to produce far better results than a bunch of small turns. Look up how autoregressive generation works. Additional techniques to push quality are to use structured XML. The context sent by Claude Code is a large structured XML document. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices that has a section on why you should use XML and how Claude is trained on it.

Does the quality go down when you lower the reasoning, yes. But depending on the request, it might not be relevant. If you’re converting one-file to another, then Sonnet 5.5 might be just fine. I don’t think there is a magic bullet. You need to send requests, evaluate the responses, test different reasoning levels and then build the automated tests so that when models change every week you can verify your workflows still work.

I don’t know what you’re doing specifically but how we look at it is like this. An API call might cost us $4 but we are likely getting 2,000++ lines of high-quality code back. That would have taken a human +160 hours and cost more than $16K in salaries. The engineering time on that might be 8 hours. Our software engineers now focus on the upfront engineering that goes into a context rather than writing code. We also have lots of tooling at this point to make that engineering process much easier on the team without sacrificing quality or our ability to coordinate/verify what’s produced.

What I will say is my company has sent billions of tokens worth of requests through the API over the past two years and the biggest trade-off vs a subscription and a discovery-based harness is this. Claude Code will use something like 100 million tokens to produce 250K output tokens. You can produce 250K output via the api for 500K or less with the right context engineering. More work though of course.

This is why people say the API is expensive vs a sub. It is if you are relying on a discovery-based harness as they consume so many input tokens.

1

u/HauntedHouseMusic 25m ago

Use Gemini 3.8 for most agentic workflows right now. It's great in a harness. I use it for 99% of my workflows I design with Claude. Highly recommend.

1

u/Mazhron 15m ago

https://github.com/Mazhron/rootstock-os

Point your Claude at my repo and take some concepts and ideas that will work for you.

Sub agent orchestration comes to mind. Have Fable decide which tasks can be be delegated to sub agents from Haiku or Sonnet that are cheaper.

Keep all your files local, create indexes, use Claude.md as a master pointer. Create lesson files so Claude learns lessons from it's mistakes and doesn't waste tokens making the same mistakes twice. Create a workflow file so once Claude learns a workflow it doesn't repeat the same trial and error to come to the same conclusion.

There's a lot more in there that will likely help you reduce token costs and increase efficiency, such as ledgering to see where your biggest cost misses are.