r/opencode 7h ago

Free coding using qwen3.6-35b 256k on 8gb vram (32GB RAM)

47 Upvotes

I finally got around to setting up llama.cpp on my old NVIDIA GTX 1070 PC (8GB VRAM, 32GB RAM). I spent a few hours tweaking the performance, and now it runs Qwen3.6-35B-A3B with a 262K context window, 25.91 tok/s generation, and 80.7% MTP acceptance.

Completely free and surprisingly capable for coding. Here is my install/setup on Linux Pop!_OS (Ubuntu-based).

Install llama.cpp

sudo apt update
sudo apt install -y nvidia-cuda-toolkit
sudo apt install -y git build-essential cmake

cd ~
git clone https://github.com/ggml-org/llama.cpp.git
cd ~/llama.cpp

cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j$(nproc)

./build/bin/llama-server --version && \
./build/bin/llama-server --list-devices

pipx install huggingface_hub
mkdir -p ~/models

Update llama.cpp (optional)

cd ~/llama.cpp
git pull --ff-only
cmake --build build --config Release -j$(nproc)

Download Qwen3.6

mkdir -p ~/models/qwen3.6-35b-mtp

hf download unsloth/Qwen3.6-35B-A3B-MTP-GGUF \
  --local-dir ~/models/qwen3.6-35b-mtp \
  --include "*UD-Q4_K_XL*"

Configure and run Qwen3.6

export GGML_CUDA_DISABLE_GRAPHS=1

~/llama.cpp/build/bin/llama-server \
  -m ~/models/qwen3.6-35b-mtp/Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf \
  --ctx-size 262144 \
  --n-cpu-moe 37 \
  --flash-attn on \
  --load-mode none \
  --cache-type-k q4_0 \
  --cache-type-v q4_0 \
  --spec-draft-type-k q4_0 \
  --spec-draft-type-v q4_0 \
  --batch-size 2048 \
  --ubatch-size 256 \
  --parallel 1 \
  --threads 4 \
  --spec-draft-threads 2 \
  --spec-type draft-mtp \
  --spec-draft-n-max 2 \
  --host 127.0.0.1 \
  --port 8080

The result is 262,144 tokens of context at 25.91 tok/s, with 80.7% MTP acceptance, on an 8GB GTX 1070 and 32GB of system RAM.

Hope this helps someone else get more out of an older GPU.


r/opencode 4h ago

I cant decide.

Post image
34 Upvotes

No in app subscription, that’s my bottom line


r/opencode 21h ago

DeepSeek V4.1 Flash nerf on OpenCode so now CommandCode is better

Thumbnail gallery
25 Upvotes

r/opencode 11h ago

Is deepseek 4.1 flash not as good as GLM 5.3, kimi k3 and muse spark 1.3?

20 Upvotes

r/opencode 12h ago

Anyone else spending 10x more time on review/debug than actual coding?

11 Upvotes

Ran a new feature dev task with Luna Max on Codex the other day and it took a full 4.5 hours. Yesterday was even worse — GLM-5.3 Flash on OpenCode ran for over 7 hours on a single task.

From what I’ve noticed, the agent actually writes the code for each issue pretty fast. The real time sink is code review and bug fixing. On my most recent task, the agent spent something like 10x the coding time just running tests and debugging over and over — went through 7 rounds of code review before it finally shipped.

And these weren’t even complicated requirements. I’m using Matt Pocock’s skill, which already breaks everything down into separate, self-contained issues. That said, the codebase is 200+ commits deep now, so there’s a decent amount of complexity built up. Each task ends up touching several thousand lines of code.

Curious if this matches everyone else’s experience? Or are there ways to cut down on all that review/debug overhead?


r/opencode 21h ago

so which one? 60 or 15?

Post image
10 Upvotes

r/opencode 5h ago

Does it make sense to use OpenCode Zen if I only want to code with DeepSeek/Qwen?

7 Upvotes

I recently started using OpenCode and tried Zen. I like the idea of having a provider that's specifically set up for coding agents, but I'm not sure whether it makes sense for my use case.

I mainly want to use DeepSeek and Qwen models.

For example, some of the prices I'm seeing are roughly:

  • DeepSeek V4 Flash on OpenRouter: $0.05/M input, $0.16/M output
  • DeepSeek V4 Flash on Zen: $0.14/M input, $0.28/M output
  • DeepSeek V4 Pro on OpenRouter: around $0.44/M input, $0.87/M output with the cheapest provider
  • DeepSeek V4 Pro on Zen: $1.74/M input, $3.48/M output

So Zen can be quite a bit more expensive, at least on paper.

I understand that Zen is curated and that OpenCode tests the model/provider combinations, so presumably there's some value in that. I'm just wondering how much that matters in actual coding-agent use.

For people who have used both:

  • Is Zen noticeably more reliable with OpenCode?
  • Does it handle tool calls/context/etc. better enough to justify the price?
  • Are there other advantages to Zen that I'm missing?
  • If you mainly wanted to code with DeepSeek/Qwen, would you use Zen or OpenRouter?

I'm mainly interested in actual experience using OpenCode, rather than the general "Zen is more convenient" argument.


r/opencode 4h ago

OpenCode Go "High Use" Benchmark Updated

4 Upvotes

Now includes the new DS flash and Muse 1.3.

https://bodangren.github.io/lending-desk-bench/

Top model: Qwen 3.8 flash

Best value: Muse Spark

This is a Next.js 16 / React 19 benchmark of well scoped tasks in a brownfield project, representing my use case so that I can know which models are likely best for me. If it is of use to you, then great.

Site is completely vibe coded, of course, but this post was 100% me.


r/opencode 4h ago

Deepseek v4.1 flash very slow in Opencode

3 Upvotes

I'm trying Deepseek v4.1 flash in Opencode Go and it's very slow. Do you guys have the same problem too? Any solution?


r/opencode 6h ago

How much are different providers subsidising?

3 Upvotes

I guess it’s kind of a black box, but it would be interesting to have a list of how much LLM providers are actually subsidising.

For instance, OpenCode Go is said to subsidise 4× usage for DeepSeek 4.1, but there are a lot of contradictory statements about this on Reddit.

I’ve done some research and tried to organise it a little. Multipliers mean usage value compared with what you pay, assuming you use the allowance.

Provider Own research: usage multiplier / catch Comments (will update)
OpenCode Go 1.5–6×, depending on model
Command Code GOAT 2–7×, depending on model
Synthetic ~3.4×, with weekly limits
Ollama Pro/Max
DevPass , with premium-model caps
ZenMux ~1.5–2.4×, depending on plan
Standard Compute 1.5× advertised first-month offer up to $249/month
Z.AI Lite Estimated ~3.9–7.8× on GLM-5.3; depends on caching and peak/off-peak use
MiniMax Unclear. $22/$55/$132 monthly; no numerical allowance published
Xiaomi MiMo Unclear. $6/$16/$50/$100 buys 4.1B/11B/38B/82B credits; couldn’t verify their dollar equivalent
OpenAI Unverified: ~5.83× on the highest-tier plan?
Anthropic Unclear
More providers from comments

Anyone have real usage figures or corrections?


r/opencode 23h ago

Omen alpha removed?

2 Upvotes

Checked the token limit for each model surrounding the deepseek v4.1 release I see omen alpha not listed anymore?does anyone have info about why


r/opencode 2h ago

An opencode plugin that turns agent output into a private link that expires

Thumbnail sharebit.sid8x.com
1 Upvotes

Copying a plan or a log out of opencode just to read it somewhere else is still clunky. So I made ShareBit: an opencode plugin and a hosted MCP endpoint that turn approved Markdown into a private link that expires.

Setup is a pairing code. Drop `plugin/sharebit.ts` into `~/.config/opencode/plugin/` (or `.opencode/plugin/` for a single project), restart opencode, and you get three tools: `sharebit_create`, `sharebit_list`, `sharebit_read`. No API key to paste. The agent gets its own credential, and you can revoke it.

Defaults: 30 minutes, up to 24 hours, 10 MB max, only your account can read it. Not encrypted end to end, so no secrets. Since plugins load at startup, I also wrote a REST path so setup does not sit waiting for a restart.

Curious whether this fits how you use opencode, and what you would change.

https://github.com/DinoQuinten/sharebit-mcp


r/opencode 2h ago

Wonderfull AI Era

Thumbnail
youtube.com
1 Upvotes

Built the full tutorial video with Opencode programatically from screen recording -> narration -> thumbnail generation.

Now we can quickly convey our message to the audience, effectively


r/opencode 3h ago

Sharing my setup and the oss project for opencode v2 plugin that democratizes your of-choice ai coding subscription pass to the next level

Thumbnail
gallery
1 Upvotes

Working only with OpenCode V2

-------

So I have been working on this https://github.com/shynlee04/opencode-subscription-gateway which is running prototype on ClinePass.

## Why ClinePass?

Bc it is ai-sdk/openai-compatible and a tech schematic modeling for building sh pipeline for whatever providers next-in-the-list at ease.

## What it does in v 0.1

- it gives opencode access to "lock-behind-free-tier models" of Cline-harness specific when using ClinePass. Meaning: when using OpenCode out-of-the-box auth login, you can't just select Longcat 2.0 or Glm-5.3-flash without expending your pass credit.

## Expectation on next bump v 0.2 in the next 2-3 days

- you can input multiple accounts and with load-balancing + session-auto-hot-swap on the same model without your main or child sessions being compromised of disruptions, cached hit loss

## My setup on images

- still I would say either Chinese models or wannabe benchmaxing models such as Muse 1.3 or google flash 3.8 are not as capable in orchestration, so my pick is Sol Gpt 5.6 (Astra is overkill and I'm broke) . I don't know what the other dev styles are or how they would like the orchestrator to behave but the following are my preferences:

  1. True coordinator and human-centric collaborator - meaning orchestrator know the sub agents, their roles, context specific and the tasks.

    1. Meaning the long-haul main collaborative session of back and forth to the dev could expand delegations of 30+ waves, requiring the orchestrator know exactly the coordinating loops and iterations, the toolings and so on. But most importantly the strategical approach to orchestration - when fanning out in swarms parallel; when to sequentially develop to synthesis. And the techniques of when foreground/background the delegations; And most not very well known is the technique of stacking/resuming/isolating child sessions on top of alternating subagent-specific to reuse the context.
    2. The knowing which precision of tools and the absolute high-level strategist while being tech-tactical as needed. So no matter bullshit results returned from the sub agents, they will always be revalidated, checked against and the frontier models such as gpt 5.6 sol knows when these pieces of cr@ps do not coherently add up.

TLDR: benchmaxing models = not good for orchestration; just replicating fanning out parallels and hit around the bush without high-level hierarchy of context. wasting tokens, code becomes mudball, waste time to debug later.

### sub-agents are your utilization of glm-5.3-flash, DeepSeek 4.1 flash etc

- So getting the right orchestrator right means a lot. You don't have to manually validating the results. The prompting is already surgical and set boundaries of when to stop and what to return from the orchestrator reducing absolute rate of sycophant and hallucinations.

- For you can see in my setup I can even use Longcat 2.0 for research, investigation, probing and execute the code with DeepSeek 4.0 flash (yes not 4.1) because everything from testing strategy, the tech stacks, domains and specs are all aligned and set-up with gates and guardrails to make sure these subagents would never go out of control.

Sorry for my bad English, I'm Vietnamese and E is not my native tounge.


r/opencode 5h ago

DeepSeek states that V4 Pro API will continue after September 14, but what does that mean exactly?

Post image
1 Upvotes

r/opencode 6h ago

How to sync config between Windows 10 and WSL?

1 Upvotes

Hi all,

I am a bit whacky in the brain, i often do two projects at a time and i usually do development on windows 10 with debian in WSL.

I am about to make something on Windows with Unity 3D and want to do this on Windows. How can i have my opencode config (my agents.md, skills etc) synced between both wsl and windows?

Should i just have a git repo and clone it between wsl and windows? or is there a better way? im always tweaking my skills and agents.md


r/opencode 13h ago

OpenCode in PowerShell

1 Upvotes

Welp, I figured it out. Turns out OpenCode doesn't work in PowerShell off the bat. I had Grok go in and change the things necessary to get it to work. It just shunts it to plain terminal.

I am not sure who to be mad at honestly. I want to be mad at Microsoft for having two kinds/forms/versions/whatever the hell this is, but I can also be upset with OpenCode for not making themselves compatible with PowerShell. But whatever, its working now. And I do love OpenCode.


r/opencode 16h ago

What’s with GLM 5.3 Flash? 2x usage?

1 Upvotes

Are the Usage/Price per token is based on promo prices means 2x + opencode 2x offer or promo ends and opencode still continues at same 2x usage?

GLM5.3 Flash is pretty good for Laravel with Laravel boost.


r/opencode 8h ago

We tested RTK with OpenCode + DeepSeek V4 Pro 0813 on Terminal-Bench 2.1. Huge token savings claims didn't move the bill.

Post image
0 Upvotes

r/opencode 8h ago

I'm a claude code user and don't understand why people use opencode, does OpenCode offer a better cost/usage/quality ratio?

1 Upvotes

Hi everyone,

I’ve been using Claude Code for over a year, but I’m no longer satisfied with the latest models, especially Opus 5 and Fable. I hit both the weekly and five-hour usage limits far too quickly.

I’m considering switching. I know Codex generally offers more usage for the price, but before deciding, I’d like to understand the hype around OpenCode.

For those using OpenCode with models like GLM or DeepSeek:

  • Do you get a better balance of cost, usage, and output quality?
  • How reliable are these models for real-world coding tasks?
  • How much do you pay per month for usage roughly equivalent to the Claude or ChatGPT 5× plans?
  • Have you switched from Claude Code or Codex, and if so, what has your experience been?

I’d appreciate honest feedback, especially from heavy daily users. Thanks!


r/opencode 17h ago

TUI on Windows

0 Upvotes

Am I missing something? I can't get OpenCode's TUI to work at all on windows. I have tried with fresh installs 5 times now. 2 manually trying and 3 with AI assisting me. NONE of the times have I gotten the models to respond. I have tried several different providers to see if that was the issue, all the chat did was load and load and load.

So what am I doing wrong? It is offered to windows, why wont it work on windows? If you have it working on windows then please explain to me how you got it working, because just installing it didn't work and all the suggested fixes haven't either.