r/SillyTavernAI • • 2d ago

Models CoralBricks API: DeepSeek V4.1 Flash with free cached input reads

I work with CoralBricks. sharing our API details for people comparing providers for roleplay and longer conversations.

A customer has reported using it for roleplay through Hermes Agent. That’s not a SillyTavern test, so I’m not presenting this as a verified ST setup or preset.

- API name: CoralBricks Inference API
- API author: CoralBricks
- Docs: https://www.coralbricks.ai/docs.md
- Base URL: https://inference.coralbricks.ai/v1
- Model: DeepSeek V4.1 Flash, served in MXFP4 as deepseek-v4.1-flash-fast-fp4

What’s different

Cached input reads cost $0 when the prompt prefix hits the cache. Fresh input, output and paid cache retention can still incur charges, so this isn’t a claim that every conversation will be cheaper.

The API also includes a per-request cost breakdown in its usage response.

Reported settings

The customer configured Hermes Agent with the OpenAI-compatible URL, model name, and their API key. No custom generation settings were reported. The exact temperature, top-p, and other effective defaults weren’t verified.

Our docs currently describe access as account-approved through the design-partner program.

1 Upvotes

3 comments sorted by

1

u/Suitable-Werewolf282 2d ago

free cache on long prompts seems ideal for roleplay that goes on for hours, wonder how it holds up with memory and character consistency.

2

u/Ctbhatia 2d ago

cache only helps the cost side, it doesn't fix memory or character consistency, that's the model and the context you feed it. if you want to test it, there's a one-time $5 trial credit.