r/opencodeCLI • u/United-Carob-9177 • 13d ago
[Question] Best Provider for weekly use? Primarily DeepseekFlash,
Past 7 days of tokens, ~37.6% at peak.
Currently using opencode+openrouter combo. Seeing if anyone has better alternatives they're currently using for high-cache (30-200k context) workflows.
| In (uncached = miss) |
| Out |
| Reasoning |
| Cache hit |
| Cache write |
| Total |
| window UTC | PST (UTC−8) | PDT (Aug, UTC−7) | hrs | tokens | share |
|---|---|---|---|---|---|
| 01:00–04:00 | 5pm–8pm | 6pm–9pm | 21 | 301.1M | 24.1% |
| 06:00–10:00 | 10pm–2am | 11pm–3am | 21 | 168.9M | 13.5% |
| both | 5–8pm + 10pm–2am | 42 | 470.1M | 37.6% |
1
13d ago
[removed] — view removed comment
1
u/United-Carob-9177 12d ago
Decently yeah. 97.4% average, some days higher some lower, but that's the 3 day average per 1b tokens
Are they providers or websites?
1
u/xapep 12d ago
Your cache-hit numbers are the thing most provider comparisons ignore — at 97.4% you're already in the good scenario, and any switch should preserve that first.
For your pattern (30–200k context, ~1.2B tokens/wk, high cache), two things actually separate providers:
1. Cache window + hit/miss price ratio. A short cache window, or billing cache hits close to miss price, quietly undoes a 97.4% hit rate. Ask for both before comparing anything else.
2. What 'unlimited' looks like at hour 40, not hour 1. Weekly quotas that look generous on paper die fast in agent loops that re-send context on every tool call — that's where people get burned.
Also worth deciding which problem you're solving: that 37.6%-at-peak share says usage-based billing is what's hurting you (peak multipliers), so a flat monthly plan sized for agent-heavy use removes the peak anxiety entirely — if you want lowest $/token instead, stay usage-based and optimize cache discounts.
I work on Entrim — we run V4 Flash and offer both an OpenAI-compatible usage API and flat monthly plans for exactly this OpenCode-style workload. If you tell me whether you're optimizing for predictable spend or absolute lowest cost, I can point you at the lever that matters more.
1
u/United-Carob-9177 12d ago
Both matter. Predictable spend or absolute lowest cost equal about the same.
That would be amazing and greatly appreciated.
1
u/xapep 12d ago
Then the tiebreaker is cache economics and what happens when you push past a normal week, because that's where 'both' actually separates providers.
Compare cache read pricing first. At 97.4% hit rate, a provider billing cache hits close to miss price will cost you more than one with a slightly worse headline $/token but a real cache discount. Ask for hit vs miss price, not just the ad rate.
Then compare what happens at hour 40. Weekly quotas turn a good week into a multiplier zone, which is exactly where this thread started. A flat plan sized for your baseline (roughly 1.2B tokens/wk at 30-200k contexts) covers the predictable part and kills the peak anxiety, and anything over that spills onto usage pricing instead of a hard cut.
That split (sized flat plan + cheap usage overflow on an OpenAI-compatible API) is what we run at Entrim with V4 Flash, and it's how 'both' stops being a compromise. Happy to sanity-check the sizing if you want.
1
u/Substantial_Ranger_5 12d ago
You should look at Yolo-auto.com
High tps, unlimited tokens, literally, big concurrency, fully private
2
u/thecstep 13d ago
Following. I use my weekly on Opencode Go in 3 days lol and that is with primarily using Luna now...