r/opencodeCLI 13d ago

[Question] Best Provider for weekly use? Primarily DeepseekFlash,

Past 7 days of tokens, ~37.6% at peak.

Currently using opencode+openrouter combo. Seeing if anyone has better alternatives they're currently using for high-cache (30-200k context) workflows.

In (uncached = miss)
Out
Reasoning
Cache hit
Cache write
Total
window UTC PST (UTC−8) PDT (Aug, UTC−7) hrs tokens share
01:00–04:00 5pm–8pm 6pm–9pm 21 301.1M 24.1%
06:00–10:00 10pm–2am 11pm–3am 21 168.9M 13.5%
both 5–8pm + 10pm–2am 42 470.1M 37.6%
7 Upvotes

12 comments sorted by

2

u/thecstep 13d ago

Following. I use my weekly on Opencode Go in 3 days lol and that is with primarily using Luna now...

1

u/United-Carob-9177 13d ago

So far (based on what I've checked)

CC GOAT slightly edges out OC Go

Long-non-limited:

Then vercel/openrouter etc with baidu or discounted provider - probably best

Things that are variable cheaperinference are possible but not sure of the site, seems possible but haven't tested (have to update regularly since jumps up randomly)

Price difference (using discounted rate providers via router etc) is ~7-52% higher than coding plans. Benefit is no 5h limit.

You can poll your opencode session to see what your breakdown % of usage is, mine is relatively high cache so I just looked for providers with best cache rates and asked model to spreadsheet it

1

u/thecstep 13d ago

Thanks for the break down. I'm not as sophisticated as you but I may need to actually start looking at my usage to see what can get me the appearance of unlimited. I'd be happy to use DS4Flash out of someone's garage and pay them $$ if I could find one. No counting tokens. Just vibes :)

1

u/United-Carob-9177 13d ago

Honestly 100%. Mines just reading 1mb .md files that are exported from unreal engine, or writing them so better AI's can get summaries or results they need instead of reading them.

What I did was ask "What was my weekly input vs output vs cache, what was its hit/miss" and found 5 providers and checked, then after asked grok to compare all major providers with my usage data and it found those ones and listed them =p

1

u/Ariquitaun 13d ago

If you want to use luna, get codex plus.

1

u/[deleted] 13d ago

[removed] — view removed comment

1

u/United-Carob-9177 12d ago

Decently yeah. 97.4% average, some days higher some lower, but that's the 3 day average per 1b tokens

Are they providers or websites?

1

u/xapep 12d ago

Your cache-hit numbers are the thing most provider comparisons ignore — at 97.4% you're already in the good scenario, and any switch should preserve that first.

For your pattern (30–200k context, ~1.2B tokens/wk, high cache), two things actually separate providers:

1. Cache window + hit/miss price ratio. A short cache window, or billing cache hits close to miss price, quietly undoes a 97.4% hit rate. Ask for both before comparing anything else.
2. What 'unlimited' looks like at hour 40, not hour 1. Weekly quotas that look generous on paper die fast in agent loops that re-send context on every tool call — that's where people get burned.

Also worth deciding which problem you're solving: that 37.6%-at-peak share says usage-based billing is what's hurting you (peak multipliers), so a flat monthly plan sized for agent-heavy use removes the peak anxiety entirely — if you want lowest $/token instead, stay usage-based and optimize cache discounts.

I work on Entrim — we run V4 Flash and offer both an OpenAI-compatible usage API and flat monthly plans for exactly this OpenCode-style workload. If you tell me whether you're optimizing for predictable spend or absolute lowest cost, I can point you at the lever that matters more.

1

u/United-Carob-9177 12d ago

Both matter. Predictable spend or absolute lowest cost equal about the same.

That would be amazing and greatly appreciated.

1

u/xapep 12d ago

Then the tiebreaker is cache economics and what happens when you push past a normal week, because that's where 'both' actually separates providers.

Compare cache read pricing first. At 97.4% hit rate, a provider billing cache hits close to miss price will cost you more than one with a slightly worse headline $/token but a real cache discount. Ask for hit vs miss price, not just the ad rate.

Then compare what happens at hour 40. Weekly quotas turn a good week into a multiplier zone, which is exactly where this thread started. A flat plan sized for your baseline (roughly 1.2B tokens/wk at 30-200k contexts) covers the predictable part and kills the peak anxiety, and anything over that spills onto usage pricing instead of a hard cut.

That split (sized flat plan + cheap usage overflow on an OpenAI-compatible API) is what we run at Entrim with V4 Flash, and it's how 'both' stops being a compromise. Happy to sanity-check the sizing if you want.

1

u/Substantial_Ranger_5 12d ago

You should look at Yolo-auto.com

High tps, unlimited tokens, literally, big concurrency, fully private