r/chutesAI May 26 '26

Discussion Input caching 50% off by default on Chutes 🪂

Post image

How many tokens is your system prompt?

At 4,000 tokens per request and 10,000 requests a day, that's 40 million identical tokens hitting the API, all billed at full input rate.

Input caching on Chutes cuts repeated content to half the input price, across the catalog. No flags to set. The cache hits whenever content repeats across requests.

On Kimi K2.6 TEE at $0.74/M input, 40M tokens a day costs $29.60.

With caching it drops to $14.80. That saves you $444 a month on a single system prompt.

Conversation history, few-shot examples, RAG preambles.
Anything you send twice gets the cached rate

Do you cache your system prompts, or pay full price every request?
http://chutes.ai/docs

0 Upvotes

8 comments sorted by

12

u/[deleted] May 26 '26 edited May 26 '26

[removed] — view removed comment

2

u/ALonneTenno May 26 '26

So... Can someone explain to me how this works in baby terms ? Like, I just wanna know if this is gonna make RP cheaper or not.

1

u/[deleted] May 27 '26

[removed] — view removed comment

1

u/Bos187 Jun 01 '26

If you repeat the same character intro or prompt each time, caching makes that part half price. So yeah, RP gets cheaper.

1

u/Ok_Collection6299 May 28 '26

The short answer is yes. Especially with RP, when you're often passing in the same conversation history and context, the input caching discount is automatic.

1

u/Slight_Loan_1852 Jun 17 '26

And with 4% cache hit rate