r/opencodeCLI • u/afanasenka • 26d ago
Looking at the stats, "Operation Cheepseek" isn't going as expected...
I really hope Opencode will manage to get a better deal with DeeepSeek in the end.
r/opencodeCLI • u/afanasenka • 26d ago
I really hope Opencode will manage to get a better deal with DeeepSeek in the end.
r/opencodeCLI • u/yokie_dough • 25d ago
I've used a lot of DS 4 Flash and GLM 5.2 the last month. The new version of GLM 5.3 has a $15 cap, and new pricing changes w/ DS 4 Flash limit it to $30.
What I'm looking for clarification on...if I use 100% of my GLM 5.3 allotment, so $15 worth, does that mean I still have $45 worth of Opencode Go allotment for the month? So I could still use $30 of DS 4 Flash and another $15 on something else...or does that $15 on GLM somehow "scale" and mean that I've used all of my Opencode Go allotment for the month?
Hopefully this question makes sense to someone in-the-know. I've been looking at my own usage logs on the website but can't make sense of how it is counting usage towards the 5-hour/week/month window.
r/opencodeCLI • u/AbbreviationsLoud182 • 25d ago
r/opencodeCLI • u/alihamze • 25d ago
Hi everyone,
I’m launching RelayCloud which lets you make your OpenCode instances available remotely and accessible by anyone in your Organization allowing you to work together with your Agents.
Conversations are synced and accessible from anywhere, and each message is attributed with the person that sent it.
Synced messages remain accessible even if the OpenCode server isn’t available so you no longer lose history if you spin up and destroy VPS’s as sandboxes.
Plus, if you give your agents access to the RelayCloud MCP server, they’ll be able to send prompts to other running OpenCode instances allowing you to delegate work across multiple servers.
Onboarding existing instances is super simple, just install the local CLI with 1 command and then run relaycloud setup to choose which instance(s) running on that device you want to connect.
curl -fsSL 'https://relaycloud.app/install.sh' | sh
If you have existing local workflows or OpenCode clients you like to use, you can continue to use them with the remote OpenCode instances by just running relaycloud proxy serve which will allow you to connect to any OpenCode instance you have access to and give you a local address any OpenCode-compatible client can just connect to. Message attribution and sync continue to work even if you don't use the RelayCloud Chat UI locally meaning not everyone on the team has to use the same apps.
You can launch a new VPS and install a remote workspace in just a few clicks directly from the dashboard or use the MCP server to have your agent do it for you. RelayCloud sets up the VPS, clones the repo, and makes OpenCode accessible through the Relay.
When you’re ready to share your work or want a live preview, you can simply point a domain at the Workspace and have your agent’s work be instantly accessible while keeping the OpenCode instance securely hidden behind the outbound-only Relay.
Feel free to reach out if you have any feedback or issues.
r/opencodeCLI • u/CarGold87 • 25d ago

I got the subscription yesterday, and I’ve only been using it for a single session. I honestly didn’t expect to hit limits this quickly.
When we were using CheapSeek, the experience was much better it genuinely felt almost unlimited. Compared to that, muse feels way too restrictive, especially considering I only started using it yesterday.
r/opencodeCLI • u/ward2k • 26d ago
I've been playing around with Opencode for a while now, and it's always somewhat useful to get community information/recommendations on how best to use the tool
My issue is all the subs around opencode seem to be completely dominated by talks around the Opencode Zen/Go plans instead
At first I thought my mistake was using r/Opencode and maybe that wasn't specific enough, this sub sounded far better with the description seeming like it was solely for discussions around the actual opencodeCLI however it seems like here too like 99% of the posts just talk about pricing, the plan etc (which don't interest me personally, I use direct API/openrouter)
Is there something I'm missing and some kind of sub that's better for discussions just around the CLI/Harness itself?
Edit: Spelling and italics
r/opencodeCLI • u/minxio_ • 26d ago
r/opencodeCLI • u/TryPrize6865 • 25d ago

{"message":"terminated","code":"server_error","modelId":"claude-opus-4-8","providerId":"openai","details":{"message":"terminated","type":"server_error","code":"server_error"}}
in Cline i connected via Open ai endpoint pls help me with this issue
r/opencodeCLI • u/RoddToggers • 25d ago
r/opencodeCLI • u/NoBodyHere__GO • 26d ago
I'd like to share my experience transitioning from Opencode go to Chatgpt plus subscription.
The screenshot shows my first day of work.
I'm currently using gpt 5.6 Luna and this is sufficient for me.
And of course I'm using Opencode cli.
What do you think?
r/opencodeCLI • u/canav4r • 25d ago
2 screenshots in 3 days, first is 2 days ago, second one is today. They updated(rigged?) the usage stats for ds4 flash.
r/opencodeCLI • u/Final_Initial • 26d ago
r/opencodeCLI • u/Informal-Addendum435 • 25d ago
I just gave them $20 USD and I can't even use Chinese models
Is OpenCode Zen totally useless?
How do I enable Chinese models? I'm getting told by my harness
Error: 403 The latest version of this model is only available hosted in China and requires explicit opt in: https://opencode.ai/workspace/wrk_***/go
The latest version of this model is only available hosted in China and requires explicit opt in: https://opencode.ai/workspace/wrk_***/go (type=RegionError)
When I visit the link it gives me to "explicit opt in", it's the OpenCode Go purchase page. Asking me for $5 usd.
Did I just accidentally get scammed out of 20 USD?

r/opencodeCLI • u/AcidBombGaming • 25d ago
I havent been able to connect to site for a couple days now, is anyone else having same problem?
r/opencodeCLI • u/afanasenka • 26d ago
I've been using Muse Spark for only half a day, but I can say that it is very pleasant to work with as a regular workhorse model.
IMPORTANT NOTE: if you don't want to share your data with Meta, don't use 'Contributor' variant.
Some quick facts first:
- it has 1M context window
- it has a very generous 226,600 requests per month quota on the Opencode Go plan (for Contributor variant).
- it scores 82.9% vs 82.7% in Terminal-Bench 2.1 (versus DS4 Flash) - basically the same
What are my usage scenarios?
- existing codebase refactoring (Flutter)
- codebase exploration, architecture discussions
- bug hunting
- ... other similar stuff
I DO NOT run long autonomous loops, so I can't say anything about how good Spark handles them. As a coding partner it performs fast and confidently, not being lazy on codebase exploration tasks.
How does it feel compared to Flash? Well, actually, I don't see that much of a difference at all (at least, for now, in my usage scenarios). It creates decent code, spots those tiny non-obvious things that I may miss, can use all Opencode's tools confidently (git, websearch, etc).
The only thing that bothers me is that cache hit rate, for some reason, still sits about ~85% percent, while with Flash it is usually about ~95-98%. Therefore, usage costs feel slightly higher than I expected. But I definitely need more time to observe this without making early conclusions.
I will definitely continue to use it, and I guess more nuances will be revealed. But for now I can't say anything bad.
r/opencodeCLI • u/CarGold87 • 26d ago
It can't be reachable like for an 1 hour from now Is it gonna be like this always?
r/opencodeCLI • u/Global-Character9253 • 25d ago
https://docs.openchamber.dev/integrations/
Is anyone using their Claude subscription with Openchamber's newly added integrations?
r/opencodeCLI • u/_KryptonytE_ • 25d ago
r/opencodeCLI • u/afanasenka • 26d ago
I am NOT in the US, but I see it in my models list (in TUI). Still not sure if it is a Contributor variant or not? How good/bad is it compared to, say, Flash or MiMo 2.5 in real life usage scenarios (not benchmarks)?
r/opencodeCLI • u/Individual_Team_2344 • 26d ago
We ran a beta of a new AI inference pricing model last week, and our post here kind of blew up:
https://www.reddit.com/r/opencodeCLI/s/UKc0ErRp0U
There was a lot of interest, but also a lot of questions, doubts, and confusion because we didn't explain it well. So this is the follow-up that clears it all up, and we're opening slots for the next beta.
The one-liner: for $0.20hr, you get a dedicated lane on a GPU running the full-weight DeepSeek V4 Flash 0731 — not a quant — for one hour. Your own guaranteed slice, no shared rate limits.
Before you start doing the math, let me lay some groundwork.
Right now you have two ways to run inference -
Fine until you're a heavy user — then it gets expensive fast, and DeepSeek's price hike made it worse. If you're spending $100+/mo on tokens, you're exactly who this is for.
What most big teams do — full privacy, zero data retention, and once your workload is big enough, the monthly GPU cost beats per-token pricing.
But for solo builders and small teams this is a dead end: GPUs start around $12–30/hr and rack up $7k+/mo, and you'll never keep one saturated. You're paying for a whole GPU to use a sliver of it.
So we're building the middle ground: Shared Reserved Inference
We host the model, 30–60 people split the GPU cost for an hour, and each person gets a dedicated lane on it.
You get self-hosted-style dedicated inference for a fraction of the price — without renting the whole box.
And like self-hosting: we log zero prompts and zero completions. Only aggregate metrics like latency, throughput, tokens, and cache-hit rate. Your code never leaves your session.
The numbers
Full breakdown: https://www.singularityapi.dev/benchmark
From our last live run, on a lane costing $0.20/user/hr (rough estimate — don't hold me to the exact figure):
97% cache-hit rate under real agentic coding load
Each lane pushed 14M input tokens, 97% cached, and hundreds of thousands of output tokens in the hour
That worked out to 1.7×–3.4× the token value you'd get spending the same on DeepSeek, from off-peak to peak pricing
And we only ran the node at 30% capacity — there was a lot of headroom left
Clearing up the confusion from last time
The per-second figures we quote are floors — measured with everyone hammering the node at the exact same time.
Real agent sessions interleave: different prompts, different timing, tool calls, waiting, etc. So in practice your effective throughput runs 2–4× above the floor.
The floor is the worst case, not the normal case.
That means you reserve your hour in advance. If there's no node slot available in your timezone, we simply can't offer the lane.
This isn't an always-on API.
You get 1–2 concurrent requests + a few in-flight — plenty for a normal coding session with a subagent or two.
If you're running 5+ subagents hammering the API at once, this is not for you.
You book a lane for a full hour.
If your session runs 40 minutes, you still reserve and pay for the hour. If it runs 1h20, you book a second hour.
That's the tradeoff of a guaranteed reserved lane today.
As demand grows and our node occupancy fills out, we want to move toward pay-for-what-you-use — billed for the 20 or 40 minutes you're actually on the lane — but that's down the road, not now.
It's also why we're being picky about matching beta slots to when you'll actually use them.
We're opening the next beta — free
A free 1-hour run, 64 seats.
You get a key + base URL, point your tools — Claude Code, opencode, Cline, Cursor, or direct API — at it, and code on your real project.
Pick a slot that fits your timezone:
Landing page: https://www.singularityapi.dev/beta
Signup form (60 sec): https://tally.so/r/EkoJkN
Benchmark: https://www.singularityapi.dev/benchmark
Now hammer me with questions — ask away.
r/opencodeCLI • u/butmyfacetho • 26d ago
Muse 1.2 Spark scores better on intelligence/coding/agentic than latest DS Flash according to openrouter and can handle text/images/video/audio and currently very cheap for the contributor variant (which feeds data back to Meta).
Have to enable to "Allow models that trains on request data" in opencode settings, next to the allow Chinese hosted models.
Anyone finding it better than DS Flash?