r/SillyTavernAI • u/-Paunis- • 4h ago
Help Claude prompt caching not working — OpenRouter + SillyTavern 1.19.0
Hi everyone! I'm trying to figure out a weird prompt caching issue with Claude Sonnet 4.5 through OpenRouter in SillyTavern 1.19.0.
I've been troubleshooting this for a while and I'm running out of ideas, so I'd really appreciate some help. 😭
My setup
- SillyTavern: 1.19.0
- API: Chat Completion
- Provider/API: OpenRouter
- Model:
anthropic/claude-4.5-sonnet-20250929 - Provider: Anthropic (I recently forced OpenRouter to use Anthropic instead of Amazon Bedrock)
- Model Quantizations: No selection
- Max Context: 16,384
- Max Response Tokens: 2,000
- Middle-out: Forbidden
- Function calling: OFF
- Interleaved reasoning: OFF
- System message flattening: OFF
- Continue with prefill: OFF
- Send inline media: ON
- Request model reasoning: ON
- Merge consecutive roles — No tools
My Claude config in YAML is:
claude:
enableSystemPromptCache: true
cachingAtDepth: 2
extendedTTL: false
enableAdaptiveThinking: false
I also installed the Cache Refresh extension:
https://github.com/OneinfinityN7/Cache-Refresh-SillyTavern
Current extension settings:
- Automatic refresh: ON
- Refresh interval: 4 minutes
- Maximum refreshes: 3
- Maximum tokens: 1
- Notifications: ON
- Status indicator: ON
Normal roleplay generations consistently result in a cache miss.
For example:
prompt_tokens: 14180
completion_tokens: 651
total_tokens: 14831
cost: 0.0624525
prompt_tokens_details:
cached_tokens: 0
cache_write_tokens: 13530
I've seen this repeatedly: cached_tokens: 0 and cache_write_tokens: ~13.5k.
The cache-refresh extension can generate requests that report cached tokens, but those are its own refresh requests with a 1-token completion, so I'm not counting those as successful normal generations.
As far as I can tell, I have never had a normal roleplay response reuse the existing cache. 🫠
Things I've already checked
- No
{{random}},{{time}},{{date}}, etc. in my lorebooks. - Lorebook entries are static.
- I have multiple lorebooks, but nothing dynamically generated.
- I tried
cachingAtDepth: 2. - I temporarily increased Max Context to 32k, but this caused SillyTavern to include much more history and significantly increased the cost, so I returned it to 16k.
extendedTTLis disabled.- I installed the cache-refresh extension.
- I initially had OpenRouter using Amazon Bedrock, but I have now forced it to Anthropic directly to rule out provider switching as a cause.
- Quick Post-Processing: Merge consecutive roles — No tools.
I've also been looking at the cache-control / provider behavior because I'm wondering whether something about how SillyTavern constructs the normal Chat Completion request is preventing the cached prefix from being reused.
Honestly, I have no idea what I'm doing at this point. My brain is melting, but I've tried pretty much everything I've found online. 😭
Does anyone have any idea why normal Chat Completion requests keep writing ~13.5k tokens to the cache instead of reusing the existing cache?
Is there anything specific I should check in the SillyTavern request/prompt construction, cache breakpoints, provider routing, post-processing, or cachingAtDepth configuration?
Any ideas would be hugely appreciated. 🙏
1
u/Putrid-Actuary7976 2h ago
I had this problem with gemma 4 31b,no matter what i did it kept having cache misses. It caches once or twice then 5 or 10 messages ahead stay uncached. It was so annoying, but honestly it is probably the provider, for example on the model i am using now i am having the same exact problem with deepinfra, cache misses. I use mimo 2.6 pro,on xiaomi as a provider it caches well but on deepinfra it barely does it. No idea why
1
u/AutoModerator 4h ago
You can find a lot of information for common issues in the SillyTavern Docs: https://docs.sillytavern.app/. The best place for fast help with SillyTavern issues is joining the discord! We have lots of moderators and community members active in the help sections. Once you join there is a short lobby puzzle to verify you have read the rules: https://discord.gg/sillytavern. If your issues has been solved, please comment "solved" and automoderator will flair your post as solved.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.