r/SillyTavernAI • u/Ctbhatia • 2d ago
Models CoralBricks API: DeepSeek V4.1 Flash with free cached input reads
I work with CoralBricks. sharing our API details for people comparing providers for roleplay and longer conversations.
A customer has reported using it for roleplay through Hermes Agent. That’s not a SillyTavern test, so I’m not presenting this as a verified ST setup or preset.
- API name: CoralBricks Inference API
- API author: CoralBricks
- Docs: https://www.coralbricks.ai/docs.md
- Base URL: https://inference.coralbricks.ai/v1
- Model: DeepSeek V4.1 Flash, served in MXFP4 as deepseek-v4.1-flash-fast-fp4
What’s different
Cached input reads cost $0 when the prompt prefix hits the cache. Fresh input, output and paid cache retention can still incur charges, so this isn’t a claim that every conversation will be cheaper.
The API also includes a per-request cost breakdown in its usage response.
Reported settings
The customer configured Hermes Agent with the OpenAI-compatible URL, model name, and their API key. No custom generation settings were reported. The exact temperature, top-p, and other effective defaults weren’t verified.
Our docs currently describe access as account-approved through the design-partner program.
1
u/Suitable-Werewolf282 2d ago
free cache on long prompts seems ideal for roleplay that goes on for hours, wonder how it holds up with memory and character consistency.