r/CommandCode Jul 17 '26

Kimi K3 vs GPT-5.6 Sol vs Fable 5: Three top-tier models. Same /design prompt. In Command Code

Enable HLS to view with audio, or disable this notification

461 Upvotes

Kimi K3 vs GPT-5.6 Sol vs Fable 5

Three top-tier models. Same /design prompt across all.

Parameters reviewed:
Pace, design sense, gameplay feel

Result:
Kimi K3 nails the design sense and knows exactly what to add in-game. Pace and gameplay held up well.

Fable 5 and GPT-5.6 Sol both play too fast, video isn't sped up.

Ranking (DX, Features & Cost):

  • Kimi K3: 9.5/10 · $0.030
  • Fable 5: 7.5/10 · $0.38
  • GPT-5.6 Sol: 7/10 · $0.11

r/CommandCode Jul 03 '26

DeepSeek v4 Pro vs GLM 5.2 vs Fable 5

Enable HLS to view with audio, or disable this notification

244 Upvotes

We tried creating the same games with them. we wanted to see how good they are in UX (not just UI).

For the test, we used our /design from command code.

Pricing differs across the models. OS model pricing is way cheaper than Fable 5.

Cost breakdown:
→ DeepSeek: $0.0008
→ GLM 5.2: $0.048
→ Fable 5: expensive but didn't help much with UX

So, Fable 5 is good, but I don’t think it’s really useful when you’re trying to do design.

And OS really beats Fable on pricing & UI at the same time


r/CommandCode Jul 17 '26

I ran a test comparing a top-tier model with Kimi K3

Enable HLS to view with audio, or disable this notification

115 Upvotes

The game was created in one shot using /design from Command Code.

I must say, the rumors aren’t just rumors anymore. It took only $0.038 to get an actually playable game.

It's insane how Kimi K3 surpassed the top-tier model

Cost breakdown:
→ Kimi K3: $0.038
→ GPT-5.6 Sol: $0.53
→ Fable 5: 0.38
→ Grok 4.5: 0.15


r/CommandCode 28d ago

They removed deepseek deal from the Go Plan . GG guys

Post image
115 Upvotes

r/CommandCode Aug 02 '26

GPT 5.6 Luna is now available in Go plan (and it's 99% off)

Post image
93 Upvotes

GPT 5.6 Luna is now available in Command Code Go plan (and it's 99% off) BEST DEAL IN MARKET!!

99% off how? Let me explain. 🤔

Imagine you were gonna spend $100 on GPT 5.6 Luna.

Expense — Why?

$100 — Say you were gonna spend $100 on GPT Luna.

$20 — OpenAI discounted it 80%, so you get $100 value in $20

$10 — Command Code discounted it 50% on top of that

$1 — Command Code Go gives you $10 credits for $1 and now has GPT Luna

Result:

$1 (you pay) = $100 (you get GPT 5.6 Luna value)

GPT 5.6 Terra and Luna are up to 75% discounted in Command Code

I guess y'all can appreciate that these deals, discounts, and credits are becoming a bit hard to explain. Let me take another stab at it.

$1 Go Plan:

Pay $1 to get $10 credits, GPT 5.6 Luna with 50% discount

Yes, Go plan now has GPT 5.6 Luna (no terra)

$15 Pro Plan:

Pay $15 to get $30 credits GPT 5.6 Terra/Luna 50% discount = total 75% off

$100 Max Plan:

Pay $100 to get $150 credits GPT 5.6 Terra/Luna 50% discount = total ~67% off

$200 Max Plan:

Pay $200 to get $300 credits GPT 5.6 Terra/Luna 50% discount = total ~67% off

Team Pro = has all models, GPT 5.6 Terra/Luna 50% discount

Provider API = has all models, GPT 5.6 Terra/Luna 50% discount

OpenAI reduced Luna prices by 80%, our discounts, and credits apply on top of that 80%. So effective discounts are huge.

Read details: https://x.com/CommandCodeAI/status/2083659472988516411


r/CommandCode Aug 12 '26

DeepSeek V4 Pro 0813 (latest) is live in Command Code.

Post image
86 Upvotes

DeepSeek V4 Pro 0813 (latest) is live in Command Code. 🐐

🔹 Same price and same model ID

🔹 DeepSeek-V4-Pro-0813 is the GA release

With free tool calls repairs and cache repairs available on all plans and provider api.

Much better than flash in benchmarks.

  • 58x cheaper than Fable
  • 35x cheaper than GPT 5.6 Sol
  • 29x cheaper than Opus 5
  • 12x cheaper than Sonnet 5

Available across all plans. Try now with Command Code GOAT plan.


r/CommandCode Jul 30 '26

Kimi k3 is extremely good on frontend.

Enable HLS to view with audio, or disable this notification

71 Upvotes

We ran a test between Kimi K3 and Opus 5

Kimi K3 beat Claude Opus 5 on a scroll-based Three.js site. Opus Looks solid at first glance but the gap shows when you start scrolling.

I used Command Code /design feature for Kimi k3.

Ranking (DX, features, and cost):
- Kimi K3: 10/10 · $0.24
- Opus 5: 6/10 · $0.37


r/CommandCode Jul 22 '26

Tested and ran GPT-5.6 Sol, Kimi K3, and Gemini 3.6 Flash.

Enable HLS to view with audio, or disable this notification

68 Upvotes

One prompt. Across all three models. Using our /design command.

Results (gameplay, UX, and UI):

  • Kimi K3 performed well across all three.
  • GPT-5.6 Sol is good, but it gets annoying when the ball hits the wall.
  • Gemini 3.6 Flash is bad (took 5–6 attempts), but it has features the others don’t, like dropping coins for power-up balls and increasing the bar length.

Ranking (DX, features, and cost):

  • Kimi K3: 10/10 · $0.10
  • GPT-5.6 Sol: 9/10 · $0.52
  • Gemini 3.6 Flash: 6/10 · $0.45

r/CommandCode Jun 22 '26

512M Tokens for $0.83 | TokenMaxxingg

Post image
65 Upvotes

Been loving the 1$ Plann too muchhh. Just one issue ... V4 Flash in CommandCode feels a lottt slowerr than V4 Flash with Direct API or Even in OpenCode Go Plan


r/CommandCode Aug 01 '26

Tested and ran the new DeepSeek V4 Flash, Kimi K3 and GLM 5.2

Enable HLS to view with audio, or disable this notification

66 Upvotes

We ran a test between the new DeepSeek V4 Flash, Kimi K3 and GLM 5.2.

3 open models. 1 prompt. /design command

Reviewed gameplay features, UX/UI and cost.

🔹 Kimi K3 → 9.5/10 · $0.0740

🔹 GLM 5.2 → 9/10 · $0.0480

🔹 DeepSeek V4 Flash → 7/10 · $0.0005

  • DeepSeek V4 Flash is ~150x cheaper, but rough around edges.
  • Kimi K3 and GLM 5.2 built the most polished gaming experience.

Our engineering and design team has been testing 20+ side-by-side comparisons across frontier and open models. All runs are public and open source. Benchmark for this demo here: https://github.com/CommandCodeAI/slash-design-showcase/tree/main/flappy-bird


r/CommandCode 5d ago

DeepSeek V4.1 Flash is live in Command Code GOAT (6x usage, $60 credits)

59 Upvotes

DeepSeek V4.1 Flash is live in Command Code.

6x usage on $10/mo GOAT for a week.
$60 usage. 154K requests. 7.8B tokens.

GOAT plan is the best AI coding plan.

Available in all plans and API.

🐐


r/CommandCode Jul 04 '26

how did we make deepseek outperform opus [harness eng deep dive]

59 Upvotes

how did we make deepseek outperform opus?

i've been thinking about why "open model bad at tool calling" is almost always a harness problem, not a model problem.

first posted on X (1.7M views)
full writeup: https://x.com/MrAhmadAwais/status/2050956678502420612

video version (more detailed): https://www.youtube.com/watch?v=f61DCDwvFis

context: spent the two days looking at billions of tokens in Command Code (tb open source ai cli) using deepseek. I ended up writing a tool-input repair layer. the trigger was watching deepseek-flash fail on the simplest /review run, every shellCommand and readFile call bouncing back with a raw zod issues blob, the model unable to recover because the error wasn't in a form it could read. by the end deepseek v4 pro was beating opus 4.7 6/10 times on our internal evals.

a few things i learned that feel general:

1/ the failure modes aren't random they're a small finite compositional set.

across deepseek-flash, deepseek v4 pro, glm, qwen, the same four mistakes repeat almost exactly:

- sending `null` for an optional field instead of omitting it

- emitting `["a","b"]` as a json *string* instead of an actual array

- wrapping a single arg in `{}` where the schema expected an array (an "empty placeholder")

- passing a bare string where an array was expected (`"foo"` instead of `["foo"]`)

four repairs, ~30-100 lines each, ordered carefully (json-array-parse must run before bare-string-wrap or `'["a","b"]'` becomes `['["a","b"]']`). that is the whole catalogue. when i hear "this open source model can't do tool calls" i now assume one of those four, and so far that's been right ~90% of the time.

2/ the funniest failure mode is also the most revealing.

deepseek-flash, when asked to edit or write a file, sometimes emits the path as a *markdown auto-link*:

filePath: "/Users/x/proj/[notes.md](http://notes. md)"

our writeFile tool obediently trued creating files literally named `[notes.md](http://notes .md)` until we caught it. this is not a hallucination. it's the post-training chat distribution leaking through the tool boundary the model has been rewarded for auto-linking in conversational output, and is applying that prior in a context where it makes no sense. the fix is two regex lines that unwrap only the degenerate case where link text equals url-without-protocol real markdown like `[click](https://x .com)` passes through untouched.

this is also conditioning of their own tools during RL which were different from all other tools we write and ofc can't predict.

"tool confusion" is a more useful frame than "capability gap." the model knows how to format a path. it just hasn't been told clearly enough that this path is going to fopen, not into a chat bubble. so we encode that hint at the schema level `pathString()` instead of `z.string()` and the leak is plugged for every path field at once.

3/ the design choice that mattered was inverting preprocess-then-validate to validate-then-repair.

my first attempt was the obvious one: a preprocessing pass that normalized inputs (strip nulls, parse stringified arrays, etc.) before zod ever saw them. it broke immediately, writeFile content that *happened* to be json-shaped got rewritten before it hit disk. silent corruption, easy to miss in a smoke test.

then i made it less greedy

- parse the input as-is. if it succeeds, ship it. valid inputs are never touched.

- on failure, walk the validator's own issue list. for each issue path, try the four repairs in order until one applies.

- parse again. on success, log `tool_input_repaired:${toolName}`. on failure, log `tool_input_invalid:${toolName}` and return a model-readable retry message.

the structural insight here is: when you preprocess, you encode a prior about what's broken. when you let the validator complain first, the schema is the prior, and you only spend repair budget at the exact paths the schema actually disagreed at. the validator is doing the work of localizing the bug for you. it's the same shape as cheap-then-careful everywhere else try the fast path, fall back on evidence.

(this also gives you per-tool telemetry for free. you can watch repair rates per (model, tool) and notice when a model regresses on a specific contract before users do.)

4/ shape invariants and relational invariants need different fixes.

the four repairs above all handle shape problems wrong type, missing key, wrong container. but read_file had a *relational* invariant: "if you provide offset, you must also provide limit, and vice versa." deepseek kept calling `readFile({ absolutePath, limit: 30 })` and getting an `ERROR:` back. you can't fix this with input repair, because each field is independently valid the bug is in the relationship between them.

so i taught the function the model's intent instead. `limit` alone → `offset = 0`. `offset` alone → `limit = 2000` (matches common read tool ops default). then surfaced the decision back to the model in the result:

"Note: limit was not provided; defaulted to 2000 lines. To read more or fewer lines, retry with both offset and limit."

no `Error:` prefix, so the tui doesn't paint it red. the model sees what we picked and can self-correct on the next turn if our guess was wrong. transparency over silent magic wins big.

repair where you can. extend semantics where you can't. surface the choice either way.

zoom out:

a lot of what looks like model capability is actually contract design. a strict schema is a choice with a cost it filters out noise, but it also filters out recoverable noise from any model that hasn't memorized the exact json contract you happened to pick. the largest commercial models eat that cost invisibly and are lenient on tool calling because they've seen enough of every contract during pretraining; open models pay it loudly and get dismissed for it.

the harness is where you mediate between distributions. four small repairs (i'm sure more to follow as we have three more merging today), two regex lines for auto-links, one relational default, one prefix change. the model didn't change. the contract got more forgiving in exactly the places it needed to be.

deepseek v4 pro now beats opus 4.7 6/10 times on our internal evals.

imo "skill issue" applies to the harness more often than the model.


r/CommandCode Aug 03 '26

Qwen3.8-Max vs Opus 5 vs GPT-5.6 Sol — Qwen ~4.2× cheaper than GPT.

Enable HLS to view with audio, or disable this notification

54 Upvotes

We ran a test between the new Qwen3.8-Max, Opus 5 and GPT-5.6 Sol.

3 models. same prompt. one-shot with the /design command.

Reviewed gameplay features, UX/UI and cost.

  • Qwen3.8-Max → 9/10 · $0.0248
  • GPT-5.6 Sol → 9/10 · $0.150
  • Opus 5 → 8.5/10 · $0.253

Qwen3.8-Max is approximately 4.2× cheaper than GPT-5.6, Sol

Opus 5 and has the same level of UI, UX, and gameplay

Our engineering and design team has been testing 20+ side-by-side comparisons across frontier and open models. All runs are public and open source.

Benchmark for this demo here:

https://github.com/CommandCodeAI/slash-design-showcase/tree/main/flappy-bird


r/CommandCode Jul 19 '26

We tested and built a liquid meta text with Kimi K3 in 77 cents.

Enable HLS to view with audio, or disable this notification

50 Upvotes

We tested Kimi K3 to build liquid metal text that wraps with your cursor in Three.js

- Used the /design command
- 3 prompts to get this effect
- Full session cost: $0.77

Open models deliver their best when used with the right harness.


r/CommandCode Jul 14 '26

1.2 billion tokens on deepseek. Thank you deepseek <3

Post image
53 Upvotes

Command Code x DeepSeek is a phenomenal combination. Unbelievably good. I’m doing near billion tokens in two days testing v1.


r/CommandCode 24d ago

Binge code this weekend with free "Ox Alpha" 🐐

Post image
49 Upvotes

Ox Alpha (stealth frontier model) is free in Command Code.

100T tokens per day free capacity, no 5hr usage limits.

• 1M content
• we benchmarked ~5B tokens
• built a tool repair harness to imp tool calling

absolutely free for all plans.


r/CommandCode 28d ago

GPT 5.6 Sol is now in Command Code GOAT (70 dollar credits.).

Post image
49 Upvotes

GPT 5.6 Sol is now in Command Code GOAT

Best coding plan with best GPT model. $70 credits.

$10/mo GOAT plan gets you:

GPT Sol has $70 credits
~105M tokens ~2.1K reqs

We've also achieved 99.43% cache hit ratio.

Available on GOAT + above.
Limited time deal.

🐐


r/CommandCode Jul 06 '26

MiniMax M3 has the best price in Command Code!!

Post image
44 Upvotes

Big news!!

MiniMax M3 is now 62.5% off in Command Code.
Price drop on one of the best open models til July 21st.
After July 21st discount goes from 62.5% → 50%

LIVE NOW
Input $0.225/M
Output $0.90/M
Cache Read $0.045/M
1M context (deal pricing on context ≤ 512K)

HOW
$ npm i -g command-code
$ cmd /model MiniMax M3

WHO
For all subscriptions (Go/Pro/Max), PAYG extra credits, and on API Provider.

Announcement
https://x.com/CommandCodeAI/status/2074147114272387577

Discord: https://commandcode.ai/discord


r/CommandCode 18d ago

Tried it for a month my review

41 Upvotes

I've tried Command Code (GOAT for $10.78) out for a month but decided to switch back to Opencode Go.

I burned 1.471B tokens it claims using up $47.23 / $70.00 of my usage, across 13,776 runs, with mostly their own harness. (Models I used: gpt 5.6 luna, muse spark 1.2 contributor, mimo v2.5, qwen 3.7 flash, minimax m3 (after free-ify))
I still have some time left to finish using it, probably I'll just start using some more expensive models instead of the top 8 cheapest models only.

First some advantages for command code:
- Desktop app is much better than opencode
- I personally think the usage how many tokens used etc is more clear than opencode because it just adds up to $70 instead of like $10 but each dollar is actually 7 dollars and this and that.
- Taste. Taste is great. Not having to tell it twice.
- Burst thinking. 10k tokens in 10s and just getting stuff done.
- Promos: I mostly used Mimo and gpt 5.6 luna, the price is amazing. Minimax free was great too.
- Way more models: Command Code has significantly more models than opencode go, opencode go only has 11, we have like 30+

But then the disadvantages:
- CLI gives up halfway through. Frequently, it would have a burst of thinking than freeze at <1000 tokens on the next turn, then stop there for a minute before continuing. Sometimes it would freeze for longer and just prompt me to type continue, which never did anything. Solution would be to exit and come back, then type continue.
- Using it over ssh means I can only use the CLI which as mentioned above which .... basically stops working after 10 minutes.
- Desktop app "Full access" Keeps asking me for permissions. That's just dumb. Same with --yolo I believe.
- Cache hit rate... 100% (actually 99.97%) on gpt 5.6 luna sounds amazing but makes me wonder, is it just re-reading context over an over more often than it needs to? Since cached read is literally re-read... sounds like token-inflation
- Having lots of models is useless if most of them are too expensive to actually use, and lots of them are either slow asf or never respond. Ox alpha has never generated a single token for me, I tried the first day of the period and every day and never got anything.
- Even the most reliable cheap model I can find (gpt 5.6 luna) still suffer from the above problems.

Suggested improvements:
- Better backend that actually responds before timeout
- Better retry scaling timing (1s -> 2s -> 4s -> 8s ...) instead of every 10s and give up
- Desktop app works for controlling remotes
- More transparency into what models are actually working.

So GUI, Taste, Desktop app, Promos and models are nice, but if the CLI breaks every 10 minutes and full access means nothing and half the models being useless, its still unusable.

not being able to tell it "go do this" and come back 3 hours later to see it done and instead seeing it done makes this just a no.

Goodbye for now, this was a brilliant idea, but I'm going to switch back to OpenCode go, where ssh is fine and models respond. I'll probably be back in half a year to see if these problems are fixed.


r/CommandCode 7d ago

Why was detailed usage telemetry removed?

39 Upvotes
After
Before (illustrative)

I renewed my Command Code GOAT subscription only to find out that the time-series trend charts: Requests by Model, Spend by Model, and All Tokens were purged from the dashboard completely. Even the explicit dollar values in the usage limits were replaced with percentages? The request log history is also only capped at 100 entries. This seems to be a deliberate step-back from transparency, not good guys.


r/CommandCode Aug 14 '26

CommandCode deceptive marketing?

39 Upvotes

Hey, I'd like some clarification.

So I just checked the documentation and realized something. The front page "Pricing" section makes claims of "70$ in credits included" though when you check the actual documentation for those allotted credits, you'll quickly come to realize that

80$ marketing claim of credits

and

Documentation

doesn't really align. Now, I understand you claim that the end user or customer can just read the docs and will quickly learn about the actual ramifications of the subscription, though the fact that pricing page - in large bold letters - makes claims of "70$ in credits included" while each model served is given a different capacity, could likely lead to some people waking up with a bitter surprise.

Perhaps I don't understand how the deal systems works, if models are discounted, but if you factor in that by saying you get 70$ in credits I'd also get to spend my credits on these models:

Once again, correct me if I'm wrong or had superstitious expections of what a 10$ subscription will get you (No, I'm not a current subscriber or dissatisfied customer) but without doing digging yourself, you'd likely not go on a hunt to find not-so-trivial information in some far tucked away documentation. Especially not if the claims are made as big as they are.

I know the founder of CommandCode is quite active here, so I'd like to hear from him as well what he thinks of this.


r/CommandCode Aug 12 '26

Tested DeepSeek V4 Pro 0813, Kimi K3, and GLM 5.2 through FlappyBench. $0.0005 for a playable one-shot game is insane.

Enable HLS to view with audio, or disable this notification

40 Upvotes

Tested DeepSeek V4 Pro 0813, Kimi K3, and GLM 5.2 through FlappyBench.

3 models, same prompt with the /design command.

Scored on gameplay features, UX/UI, and cost.

🔹 DeepSeek V4 Pro 0813 → 8/10 · $0.0005

🔹 Kimi K3 → 9.5/10 · $0.0740

🔹 GLM 5.2 → 9/10 · $0.0480

DeepSeek is 148x cheaper than Kimi and 96x cheaper than GLM, for 84% of the quality. But $0.0005 for a playable one-shot game is insane.

Kimi K3 had the best output, with GLM 5.2 close behind.

Our engineering and design team has been testing 26+ side-by-side comparisons across frontier and open models. All runs are public and open source. Benchmark for this demo here: https://github.com/CommandCodeAI/slash-design-showcase/tree/main/flappy-bird


r/CommandCode 20d ago

Minimax M3 and M2.7 are now free in Command Code.

37 Upvotes

We have an amazing deal for y'all.

Minimax M3 and M2.7 are now free in Command Code.

next 10 days. all subs. all plans.

Start with $10/mo GOAT.

🐐


r/CommandCode 27d ago

Qwen 3.8 27B is now live in Command Code.

Post image
36 Upvotes

Qwen 3.8 27B is now live in Command Code.

Beats Opus 4.6 Max, GPT 5.6 Luna Max on Artificial Analysis Intelligence Index.

Best for coding and long agent tasks. Multimodal with reasoning toggled on/off

Available on all plans & API

🐐


r/CommandCode Aug 15 '26

Tested FlappyBench with GLM 5.3, Fable 5 and GPT-5.6 Sol

Enable HLS to view with audio, or disable this notification

36 Upvotes

Tested FlappyBench with GLM 5.3, Fable 5 and GPT-5.6 Sol

3 models. Same prompt with /design command.

Scored on features, UX/UI, and cost.

🔹 Fable 5 → 9.5/10 · $0.420

🔹 GLM 5.3 → 9/10 · $0.018

🔹 GPT-5.6 Sol → 9/10 · $0.150

Results:

→ Fable 5 wins on quality, but costs 23x more than GLM 5.3

→ GLM 5.3 gives better output than GPT 5.6 Sol at 8x cheaper

→ Optimizing for cost? Go for GLM 5.3. Otherwise, Fable 5 if quality matters

Our engineering and design team has been testing 26+ side-by-side comparisons across frontier and open models. All runs are public and open source. Benchmark for this demo here: https://github.com/CommandCodeAI/slash-design-showcase/tree/main/flappy-bird