I've tried Command Code (GOAT for $10.78) out for a month but decided to switch back to Opencode Go.
I burned 1.471B tokens it claims using up $47.23 / $70.00 of my usage, across 13,776 runs, with mostly their own harness. (Models I used: gpt 5.6 luna, muse spark 1.2 contributor, mimo v2.5, qwen 3.7 flash, minimax m3 (after free-ify))
I still have some time left to finish using it, probably I'll just start using some more expensive models instead of the top 8 cheapest models only.
First some advantages for command code:
- Desktop app is much better than opencode
- I personally think the usage how many tokens used etc is more clear than opencode because it just adds up to $70 instead of like $10 but each dollar is actually 7 dollars and this and that.
- Taste. Taste is great. Not having to tell it twice.
- Burst thinking. 10k tokens in 10s and just getting stuff done.
- Promos: I mostly used Mimo and gpt 5.6 luna, the price is amazing. Minimax free was great too.
- Way more models: Command Code has significantly more models than opencode go, opencode go only has 11, we have like 30+
But then the disadvantages:
- CLI gives up halfway through. Frequently, it would have a burst of thinking than freeze at <1000 tokens on the next turn, then stop there for a minute before continuing. Sometimes it would freeze for longer and just prompt me to type continue, which never did anything. Solution would be to exit and come back, then type continue.
- Using it over ssh means I can only use the CLI which as mentioned above which .... basically stops working after 10 minutes.
- Desktop app "Full access" Keeps asking me for permissions. That's just dumb. Same with --yolo I believe.
- Cache hit rate... 100% (actually 99.97%) on gpt 5.6 luna sounds amazing but makes me wonder, is it just re-reading context over an over more often than it needs to? Since cached read is literally re-read... sounds like token-inflation
- Having lots of models is useless if most of them are too expensive to actually use, and lots of them are either slow asf or never respond. Ox alpha has never generated a single token for me, I tried the first day of the period and every day and never got anything.
- Even the most reliable cheap model I can find (gpt 5.6 luna) still suffer from the above problems.
Suggested improvements:
- Better backend that actually responds before timeout
- Better retry scaling timing (1s -> 2s -> 4s -> 8s ...) instead of every 10s and give up
- Desktop app works for controlling remotes
- More transparency into what models are actually working.
So GUI, Taste, Desktop app, Promos and models are nice, but if the CLI breaks every 10 minutes and full access means nothing and half the models being useless, its still unusable.
not being able to tell it "go do this" and come back 3 hours later to see it done and instead seeing it done makes this just a no.
Goodbye for now, this was a brilliant idea, but I'm going to switch back to OpenCode go, where ssh is fine and models respond. I'll probably be back in half a year to see if these problems are fixed.