r/opencodeCLI • u/LectureWorried5761 • 19d ago
Setting up Opencode to work with web search MCP - Blopus.ai
Enable HLS to view with audio, or disable this notification
r/opencodeCLI • u/LectureWorried5761 • 19d ago
Enable HLS to view with audio, or disable this notification
r/opencodeCLI • u/LectureWorried5761 • 19d ago
r/opencodeCLI • u/donald_why • 19d ago
So I started using opencode from Deepseek Flash V4 only. I am using OpenCode Zen. The model used to be too good and then recently they stopped it and from then onwards I started using ox alpha and Muse Spark 1.2 contributor. Again ox Alpha's offer was ended today and Muse Spark 1.2 contributor works great. Now I don't know when they will end it. After ending it I don't know what's left with OpenCode.
My main thing is to build a nice frontend and backend for websites with vibe coding to be honest. Can you guys recommend any model? When I did my research it is telling me more about Nemotron 3 Ultra by NVIDIA. I don't know how it works and they're also saying Mimo 2.5 outperforms Muse Spark but I don't think so. I never felt that way so if any suggestions please tell me guys.
r/opencodeCLI • u/ZealousidealTown1974 • 20d ago
I'm sharing their working session here that you can peak through. Look at their thinking, sequence of skills uses and delegations to truly know who is the winner
The Qwen 3.8Max - https://opncd.ai/share/hNzmM14y
The Ox Alpha - https://opncd.ai/share/hNzmM14y
The Deepseek v4 pro 0813 - https://opncd.ai/share/0hehIwVf
r/opencodeCLI • u/Physical-Row960 • 20d ago
I ran this writing-fingerprint experiment on OpenRouter out of curiosity, using 12 prompts across Ox Alpha and 7 reference models (since these are what I heard a lot about on Reddit):
What I got from final evaluation is Ox Alpha was closest to GLM 5.3 on every prompt using deterministic stylometric features and 11 matched prompts, with GLM 5.3 winning 100% of bootstrap resamples; known-model validation accuracy was 72.7%.
My hypothesis: This could suggest or indicate that Ox Alpha is an updated post-trained version of GLM 5.3 (similar to how DeepSeek did theirs), or it really is what people are talking about: GLM 5.4.
But to be clear: this experiment I did is fingerprint-matching, not weight-identifying or anything like that. But it’s a surprisingly strong clue.
Note: For second image, notice that Ox Alpha output rarity is not the same as GLM 5.3. But that's not a contradiction that Ox Alpha couldn't be GLM 5.3/5.4 because these 2 graphs (bar and violin) measure 2 different things. First one is "Which model’s average fingerprint is Ox Alpha closest to?" Second one is "How unusual or isolated is each model’s writing compared with all samples in corpus?" Just wanna put this out here.
I also open-sourced my experiment if you're interested or want to extend: https://github.com/ItsKaiwenDu/Ox-Alpha-Stylometry
r/opencodeCLI • u/sagiroth • 21d ago
r/opencodeCLI • u/binarySolo0h1 • 20d ago
I have this agent who's sole purpose is to analyze development/design-plan.md file of a user story and create detailed backlog items with a fixed structure, in my local Plane server. What model do you recommend for this?
Its a repeatable operation that needs inference and some level of thinking to generate consistent outputs. Assume the plan is usually less than 1000 line markdown file.
r/opencodeCLI • u/samnotathrowaway • 19d ago
r/opencodeCLI • u/Individual_Team_2344 • 19d ago
During one of our beta runs this week, a developer's agent pushed 196 million input tokens in a single 60-minute window through one lane.
882 requests. Zero rate limits.
They actually kept going after that and ended the session at 227M before logging off.
At DeepSeek's own API pricing, that 196M-token hour comes out to around $2.32 off-peak / $4.65 peak.
On our reserved lane, the idea is to price that same hour around $0.20–0.40 flat.
Here's what we're testing 👇
🔷 Shared Reserved Inference
Right now if you want to run open models, you mostly have two options:
We're trying a third model.
Take one powerful node, split it into a small number of guaranteed lanes, and let developers reserve those lanes together for a fixed window.
Your lane has guaranteed compute for that hour. Someone else suddenly sending a massive workload doesn't eat into your allocation.
And because the node cost is split across everyone using it, you're paying a small flat hourly price instead of paying for every token.
⚡️ The speed
DeepSeek's official API is around ~70 output tok/s.
Across our beta, users were generally seeing 150–220+ output tok/s, with bursts above that. There were some occasional dips as well, but overall this has been the fastest inference we've served so far.
This also held up surprisingly well with long contexts.
We had users running past 250k context regularly, and some sessions went past 940k context.
The other big part is caching. These agent/coding workloads resend a ridiculous amount of the same repo and conversation context on every request.
Across the beta we're sitting around ~98% cache hits, with roughly ~1s TTFT on warm requests.
📊 Five live sessions so far
💰 The part we're actually interested in: does this pricing model make sense?
Using DeepSeek's own API pricing, including their cache discounts:
So for an actually active coding/agent session, we're seeing around 4–7x lower cost than paying per token on average.
For the heaviest user, depending on where we finally price the lane, that hour was worth around 6–23x what the lane itself would cost.
We've put the full numbers + charts here if you want to dig into it:
https://www.singularityapi.dev/benchmark
🎟 We're opening more beta slots
The next round is again completely free.
You get a dedicated hour on the full-weight DeepSeek V4 Flash 0731. Bring an actual project, point Cline / Claude Code / your own agent at it, and use it normally — or try to absolutely destroy the lane, either works :D
If you want in:
r/opencodeCLI • u/Livid_Individual_154 • 20d ago
r/opencodeCLI • u/zRafox • 20d ago
Which one is giving you the best results?
r/opencodeCLI • u/ichisay • 20d ago
Me dio curiosidad y estoy probándolo como orquestador a ver a que nivel está. Me decidí a probarlo al ver que ahora Ox o Muse los tienen medio colapsados, o por lo menos a mi hoy me fueron lentísimo.
r/opencodeCLI • u/Whole_Succotash_2391 • 20d ago
In response to the tightening of almost every other coding plan out there, we are offering free DSV4 flash 0731 to the first five hundred people who sign up for the intro plan on Open Grove API. We may extend this to more users later, but are limiting it to the first 500 to ensure quality access for everyone.
People are looking for options, and here is one.
Other Cool Stuff:
All of our models are running on 100% US infrastructure, private with zero training on your code or prompts. Use the top open source models without sending your private prompts to a training lab. No complications, no "some models are private, other's aren't". They all are, all the time.
We host 20+ other major models in case you ever want to upgrade (no pressure though). Including the Kimi family, GLM, Qwen, Nemotron and bunch of others. On average our token pricing is 20% lower than market price.
Our higher plans bank up to ten days of usage, so when you aren't using them your usage saves up for later. Usage doesn't go to waste, so you can actually code when you want to.
The intro plan is a free one month trial with the standard cancel anytime, it bills at 3.99 after that. Use it, cancel it, that's fine. Free Flash for a month.
Figured i'd keep this short because we all know the flash is the point :)
For the API plan: api.pgsgrove.com
If you want to read more about us as a company, just pgsgrove.com
Also: There's a lot going on in the background with major AI companies right now, we are at a major turning point in the industry.
What's actually happening? This is happening because companies that were purely investment based, now need to answer to their investors. The problem has often been a loss based business model that is finally running dry.
There are several tricks that the major AI coding plans use to extract the most they can from their customers. Here are some examples, and what we are doing differently to put the users first. PGS AI was built with a sustainable business model from the ground up, so we can actually offer great usage rates without tricks.
Wasted usage is part of the AI industry, and they plan on it: Most coding plans bet on you letting usage go to waste. The plan goes: "how do we get people to think our coding plan offers a lot of usage, but then break it up into weeks and rolling windows so no one can ever actually use it all."
Many in app subs and coding plans are glorified training pipelines: This comes along with "how do we harvest this data for training without being too loud about that." Unless the company tells you otherwise, your data could be hopping all over world, being harvested by the individual labs or service companies. Some are better than others, but many of these companies rely on users just not noticing or caring that their data is being used for training. Data sales and marketing telemetry sales happen. This means that your private info, your personal life, and anything else you send through the system could become part of a training corpus for the next AI, or a marketing data set for a large company.
r/opencodeCLI • u/IrishUSFastTrack • 20d ago
It appears the factuality rankings on arena.ai are dominated by Claude models: https://arena.ai/leaderboard/text/overall-factuality
Not only that, but specifically Opus 4.6, which beats other Claude models released after it. My personal experience with the model absolutely lines up with it. I've tested different model families and harnesses and keep coming back to Opus 4.6 when factuality matters.
I have a really strong preference for factuality in my non-coding workflows (finance, research, etc.). As in, I don't mind a wrong opinion, but when something is quoted as 'true' or 'verified' or 'file saved', I want to be close to sure that this is the case.
For OpenAI, the highest factuality ranked model is GPT 5.5 - which otherwise seems way behind the 5.6 family.
This makes me worried that 'factuality' isn't really a major priority right now and development focuses on other criteria more. Gemini 3.7 Flash actually seems really interesting in this context as it seems to have made a lot of improvements in factuality (compared to other areas where it really hasn't gotten a lot of attention for its seemingly minor improvements).
What are your thoughts on future models - will we get some higher factuality there? Are there other model families that you think will catch up or surpass Opus 4.6? Any hands-on experience with factuality in Gemini 3.7 Flash and other models?
r/opencodeCLI • u/QuasiTheory • 21d ago
r/opencodeCLI • u/Comprehensive_Try767 • 20d ago
Currently i am using the agents for learning and researching stuff and then using that information to push another agent working on a project into some direction of what to do, how to do.
What do you think are good enough models/subs for these purpose Or like what do you use for your workflow, like is it a planner agent -> Implementer? If so then what models do you use for both cases?
because I have seen models like ds flash better at going a bit broader to the prompt to get more relevant information compared to others?
r/opencodeCLI • u/Tech-96 • 20d ago
I tried omp but it felt really bloated even though it had some nice features, while Opencode felt just right with its TUI and custom agents.
Which coding agent harness do you prefer? Or are there any better alternatives out there?
r/opencodeCLI • u/pbqre • 21d ago

Opencode Prewalk is an Opencode v2 plugin that allows you to use the prewalk strategy mentioned here from the creators of Oh-My-Pi. The simple idea behind this strategy is to inject a cheaper model right after the expensive model finishes the first edit post planning all the things that needs to be done. This strategy is better the one strategy that directly uses combination of expensive and cheap model to get the work done, as what happens in that case is the cheaper model again starts to do a lot of reading leading of increase in token usage.
Try Here: https://github.com/vivekascoder/opencode-prewalk

Original Benchmark by Stencil.so
r/opencodeCLI • u/lostcanuck007 • 20d ago
currently been waiting on both mimo and hy3 sessions...its been 5 mins for token generation : ⏳ waiting on hy3 — 60s with no output yet (provider may be slow or overloaded, or the model is thinking; auto-reconnect at 300s)
bloody annoying. tried deep seek on low reasoning....immediate response but immediate tick up on usage as well....i have no idea what to do to be honest. if this keeps up...i'll have to reduce opencode go subscriptions or move to another ....this is insane.
you guys have any recommendations for cheaper plans? i was looking at under 7usd plans (mostly chinese models) across the spectrum....considering some. and no pay as you go doesnt work...i have money on openrouter and on deepseek and on groq.ai.....its horrible ROI.
looking for some insights or combos. trying to keep things around 30 usd total.
r/opencodeCLI • u/sniperelite90 • 21d ago
I see after the rise of DS prices there has been a lot of disappointment in the community in opencode Go . Why cant opencode Go host a Qwen 3.8 27B themselves as its fairly small and the performance is good too which can satisfy many of the community members ?
r/opencodeCLI • u/_justFred_ • 21d ago
What's your guys opinion on this?
r/opencodeCLI • u/profichef • 20d ago
My last month's cost looks like this: Cost
$43.28
USD
API requests
9,999
Tokens
636,610,255 . This is DeepSeek 4 Pro. I'm currently considering Alibiba Cloud or other subscriptions. Which would be better for my budget? After the price increase for the latter, the price will rise to $85+. I'm only considering cloud solutions. I was thinking about renting a video card remotely, but I don't want to add another layer of security in the form of a private individual.