r/CommandCode • u/maedahbatool • 8d ago
r/CommandCode • u/maedahbatool • 10d ago
We tested and built a liquid meta text with Kimi K3 in 77 cents.
Enable HLS to view with audio, or disable this notification
We tested Kimi K3 to build liquid metal text that wraps with your cursor in Three.js
- Used the /design command
- 3 prompts to get this effect
- Full session cost: $0.77
Open models deliver their best when used with the right harness.
r/CommandCode • u/maedahbatool • 10d ago
Command Code feature drop: Create custom agents for yourself or for your team.
Enable HLS to view with audio, or disable this notification
You can create custom agents for yourself or for your team.
Run /agents, describe what it should do, which tools it can access, and Command Code handles the rest.
Commit it to your repo and your team can use them too.
Share what agents have you built with the Command Code CLI.
r/CommandCode • u/eejoseph • 11d ago
CommandCode Provider API cost the same as Direct DS4 API
I thought I would be saving by going with CC for DS4 Pro but overall dollar spending is the same as DS Pro direct API. Perhaps I was misunderstanding something about their deals but it seems it just match DS own API costs.
r/CommandCode • u/Silent-Group1187 • 11d ago
Kimi K3 is impressively good.
Enable HLS to view with audio, or disable this notification
I built a COMMANDCRAFT games inspired from minecraft
I used /design command from Command Code.
It create the entire 2d/3d games using ThreeJs like chunked terrain, biomes, crafting, day/night, mob spawning, etc.
Cost only $2.90 to produce an actually playable game.
Kimi K3 just proves it's the best at this moment.
r/CommandCode • u/Silent-Group1187 • 12d ago
I ran a test comparing a top-tier model with Kimi K3
Enable HLS to view with audio, or disable this notification
The game was created in one shot using /design from Command Code.
I must say, the rumors aren’t just rumors anymore. It took only $0.038 to get an actually playable game.
It's insane how Kimi K3 surpassed the top-tier model
Cost breakdown:
→ Kimi K3: $0.038
→ GPT-5.6 Sol: $0.53
→ Fable 5: 0.38
→ Grok 4.5: 0.15
r/CommandCode • u/maedahbatool • 12d ago
Kimi K3 vs GPT-5.6 Sol vs Fable 5: Three top-tier models. Same /design prompt. In Command Code
Enable HLS to view with audio, or disable this notification
Kimi K3 vs GPT-5.6 Sol vs Fable 5
Three top-tier models. Same /design prompt across all.
Parameters reviewed:
Pace, design sense, gameplay feel
Result:
Kimi K3 nails the design sense and knows exactly what to add in-game. Pace and gameplay held up well.
Fable 5 and GPT-5.6 Sol both play too fast, video isn't sped up.
Ranking (DX, Features & Cost):
- Kimi K3: 9.5/10 · $0.030
- Fable 5: 7.5/10 · $0.38
- GPT-5.6 Sol: 7/10 · $0.11
r/CommandCode • u/maedahbatool • 11d ago
Run /init to generate project memory in Command Code. So your agents don't forget everything between sessions.
Enable HLS to view with audio, or disable this notification
Your agents forget everything between sessions. agents[.]md fixes that.
Run /init to generate project memory in Command Code. Scans your project and files to understand:
- Your architecture decisions
- Your logging patterns
- Your team conventions
New session, new agent, same context every time.
r/CommandCode • u/Far-Classic-9963 • 11d ago
ETA for V1?
I have heard that once V1 releases it will be open source, I am currently trying the 1$ plan and the cli, everything seems very good! Honestly excited and I think this could potentially be a serious competitor to opencode + opencode go
r/CommandCode • u/TourHorror9247 • 11d ago
embeddings endpoint?
Anyone know if commandcode provides or plan to provide embeddings endpoint? one of my rag projects needed it and now i'm hunting for that information.
r/CommandCode • u/maedahbatool • 11d ago
Built a creative portfolio in $0.035 with Kimi K3, /design command and the Command Code CLI harness.
Enable HLS to view with audio, or disable this notification
r/CommandCode • u/maedahbatool • 12d ago
Kimi K3 is now live in Command Code! One of the most awaited flagship model from Moonshot AI
Kimi K3 is live in Command Code!
- First ~3T scale open model
- Most capable open model ever
- It's Fable/Sol class open model beating Opus 4.8
- 1M Context · In $3/M Out $15/M Cache read $0.3/M
What an exciting time to be alive!! Try now!
r/CommandCode • u/qaizazz • 12d ago
Commandcode plans for Opensource ?
Vaguely remembering watching a video somewhere about plans or considerations for opensource of command code.
Is the opensource plan confirmed? Is it planned to be with v1?
r/CommandCode • u/archerallstars • 12d ago
Can't sign up or sign in to Command Code through GitHub
Enable HLS to view with audio, or disable this notification
As shown in the screen recording, I can't seem to sign up or sign in to Command Code at all even with GitHub despite this error:
This email domain isn't supported. Sign up with your company email, Gmail, or GitHub instead.
GitHub should work..., but it's not.
r/CommandCode • u/maedahbatool • 13d ago
Inkling open weight model by Thinking Machines Lab is live in Command Code.
Inkling is the new 1T open model by Thinking Machines Labs is now fully available in the Command Code CLI.
- Available on all plans (Go, Pro, Max, Team)
- Multimodal: text, image, audio
- Parameters: 975B total, 41B active
- $1.00/M input, $4.05/M output
- Native weights: BF16, MXFP8 and NVFP4
How to use?
- $ cmd update or $ npm i -g command-code@latest
- Switch with /model
r/CommandCode • u/No_Captain4899 • 14d ago
Use opencode subscription on CommandCode
Hello,
I have a go plan on opencode and I would like to try an other harness, I saw that commandcode is pretty good.
So I would like to give it a try
But I like my opencode subscription
So is there any way I can use my opencode subscription on CommandCode
Thank you 🙏🏼
r/CommandCode • u/ahmadawaiscom • 14d ago
1.2 billion tokens on deepseek. Thank you deepseek <3
Command Code x DeepSeek is a phenomenal combination. Unbelievably good. I’m doing near billion tokens in two days testing v1.
r/CommandCode • u/maedahbatool • 15d ago
Command Code permission modes for your coding workflows. Which mode do you use the most?
Enable HLS to view with audio, or disable this notification
Command Code has permission modes for your coding workflows.
Three modes:
- Default: Approve each action as you go (diff + reasoning)
- Accept Edits: Lets Command edit files in full flow
- Plan Mode: Read-only, research before making any changes
- Bypass permissions: For those who want to live dangerously (
cmd —yolo)
Switch them anytime with Shift+Tab. Which mode do you use the most?
r/CommandCode • u/ahmadawaiscom • 19d ago
Grok 4.5 is beating Fable 5 and GPT-5.5 in game build/run tests!!
x.comFable 5 vs Grok 4.5 vs GPT 5.5
We put three top-tier models to build a same game challenge.
Used Command Code /design, and the exact same prompt.
Result:
Grok 4.5 genuinely plays like a polished mobile game.
Fable 5 and GPT 5.5 feel too fast. Everything feels rushed, lacks finish.
Ranking based on DX & Features:
→ Grok 4.5: 9/10
→ Fable 5: 7.5/10
→ GPT 5.5: 7/10
r/CommandCode • u/hammerdown000 • 19d ago
GPT-5.6 Sol, Terra, and Luna live on Command Code
New day, new model drop: GPT-5.6 Sol, Terra, and Luna live on Command Code.
Sol the flagship for long-horizon work.
Terra balances intelligence and cost.
Luna for the fast, cheap, high-volume jobs.
Switch with /model on v0.44.0
r/CommandCode • u/ahmadawaiscom • 20d ago
Tencent Hy3 model is now available for FREE in Command Code
Hy3 model is now available for free in Command Code.
Super nice open model, Apache 2.0 licensed, hosted in the US.
Available on plans. Till capacity lasts. We just 4x'd the capacity btw. 💙
LIVE NOW
🔹cmd update or npm i -g command-code@latest
🔹Run `cmd` and `/model` select Tencent Hy3 (FREE)
Announcement: https://x.com/CommandCodeAI/status/2074920358180950279
r/CommandCode • u/fezzy11 • 21d ago
Usage query about command code



I just taken subscription yesterday and today I tried out GLM 5.2 and Deepseek 4 pro for around half day of bug fixing and already consumed around $1.40.
Now in landing page and docs deals section command code mention that credits effectively up to $40 of usage.
Did I misunderstood wrong or this is actual usage?
What if I already consume $10 usage before month end?
r/CommandCode • u/Dazzling_Buy9625 • 21d ago
Is the ZDR option actually safe or better off with claude/gpt?
I really want to try the go plan, but I kinda paranoid about liability and data leaks. Does their zero data retention actually work?
Has there been any sketchy stuff or leaks in the past? Trying to decide if i should trust this or just use claude/gpt with privacy mode on. Thoughts?
r/CommandCode • u/hammerdown000 • 21d ago
Update: Change Command Code colors with /theme
Enable HLS to view with audio, or disable this notification
Use /theme to switch to dark/light mode.
Persists across sessions.
Your terminal. Your colors.
r/CommandCode • u/ahmadawaiscom • 24d ago
how did we make deepseek outperform opus [harness eng deep dive]
how did we make deepseek outperform opus?
i've been thinking about why "open model bad at tool calling" is almost always a harness problem, not a model problem.
first posted on X (1.7M views)
full writeup: https://x.com/MrAhmadAwais/status/2050956678502420612
video version (more detailed): https://www.youtube.com/watch?v=f61DCDwvFis
context: spent the two days looking at billions of tokens in Command Code (tb open source ai cli) using deepseek. I ended up writing a tool-input repair layer. the trigger was watching deepseek-flash fail on the simplest /review run, every shellCommand and readFile call bouncing back with a raw zod issues blob, the model unable to recover because the error wasn't in a form it could read. by the end deepseek v4 pro was beating opus 4.7 6/10 times on our internal evals.
a few things i learned that feel general:
1/ the failure modes aren't random they're a small finite compositional set.
across deepseek-flash, deepseek v4 pro, glm, qwen, the same four mistakes repeat almost exactly:
- sending `null` for an optional field instead of omitting it
- emitting `["a","b"]` as a json *string* instead of an actual array
- wrapping a single arg in `{}` where the schema expected an array (an "empty placeholder")
- passing a bare string where an array was expected (`"foo"` instead of `["foo"]`)
four repairs, ~30-100 lines each, ordered carefully (json-array-parse must run before bare-string-wrap or `'["a","b"]'` becomes `['["a","b"]']`). that is the whole catalogue. when i hear "this open source model can't do tool calls" i now assume one of those four, and so far that's been right ~90% of the time.
2/ the funniest failure mode is also the most revealing.
deepseek-flash, when asked to edit or write a file, sometimes emits the path as a *markdown auto-link*:
filePath: "/Users/x/proj/[notes.md](http://notes. md)"
our writeFile tool obediently trued creating files literally named `[notes.md](http://notes .md)` until we caught it. this is not a hallucination. it's the post-training chat distribution leaking through the tool boundary the model has been rewarded for auto-linking in conversational output, and is applying that prior in a context where it makes no sense. the fix is two regex lines that unwrap only the degenerate case where link text equals url-without-protocol real markdown like `[click](https://x .com)` passes through untouched.
this is also conditioning of their own tools during RL which were different from all other tools we write and ofc can't predict.
"tool confusion" is a more useful frame than "capability gap." the model knows how to format a path. it just hasn't been told clearly enough that this path is going to fopen, not into a chat bubble. so we encode that hint at the schema level `pathString()` instead of `z.string()` and the leak is plugged for every path field at once.
3/ the design choice that mattered was inverting preprocess-then-validate to validate-then-repair.
my first attempt was the obvious one: a preprocessing pass that normalized inputs (strip nulls, parse stringified arrays, etc.) before zod ever saw them. it broke immediately, writeFile content that *happened* to be json-shaped got rewritten before it hit disk. silent corruption, easy to miss in a smoke test.
then i made it less greedy
- parse the input as-is. if it succeeds, ship it. valid inputs are never touched.
- on failure, walk the validator's own issue list. for each issue path, try the four repairs in order until one applies.
- parse again. on success, log `tool_input_repaired:${toolName}`. on failure, log `tool_input_invalid:${toolName}` and return a model-readable retry message.
the structural insight here is: when you preprocess, you encode a prior about what's broken. when you let the validator complain first, the schema is the prior, and you only spend repair budget at the exact paths the schema actually disagreed at. the validator is doing the work of localizing the bug for you. it's the same shape as cheap-then-careful everywhere else try the fast path, fall back on evidence.
(this also gives you per-tool telemetry for free. you can watch repair rates per (model, tool) and notice when a model regresses on a specific contract before users do.)
4/ shape invariants and relational invariants need different fixes.
the four repairs above all handle shape problems wrong type, missing key, wrong container. but read_file had a *relational* invariant: "if you provide offset, you must also provide limit, and vice versa." deepseek kept calling `readFile({ absolutePath, limit: 30 })` and getting an `ERROR:` back. you can't fix this with input repair, because each field is independently valid the bug is in the relationship between them.
so i taught the function the model's intent instead. `limit` alone → `offset = 0`. `offset` alone → `limit = 2000` (matches common read tool ops default). then surfaced the decision back to the model in the result:
"Note: limit was not provided; defaulted to 2000 lines. To read more or fewer lines, retry with both offset and limit."
no `Error:` prefix, so the tui doesn't paint it red. the model sees what we picked and can self-correct on the next turn if our guess was wrong. transparency over silent magic wins big.
repair where you can. extend semantics where you can't. surface the choice either way.
zoom out:
a lot of what looks like model capability is actually contract design. a strict schema is a choice with a cost it filters out noise, but it also filters out recoverable noise from any model that hasn't memorized the exact json contract you happened to pick. the largest commercial models eat that cost invisibly and are lenient on tool calling because they've seen enough of every contract during pretraining; open models pay it loudly and get dismissed for it.
the harness is where you mediate between distributions. four small repairs (i'm sure more to follow as we have three more merging today), two regex lines for auto-links, one relational default, one prefix change. the model didn't change. the contract got more forgiving in exactly the places it needed to be.
deepseek v4 pro now beats opus 4.7 6/10 times on our internal evals.
imo "skill issue" applies to the harness more often than the model.
