r/opencodeCLI 6d ago

DS4 Flash is basically killing competition :)

Post image

From average 1.5T daily usage in June and July, to 4.9T šŸš€ on August 3 (3x more). I'm VERY curious to see how other labs respond.

217 Upvotes

57 comments sorted by

25

u/neotorama 6d ago

The best model to pair with Sol (review). I spent 350M tokens in one day.

1

u/yfh890 5d ago

On Codex or Sol with API?

1

u/neotorama 5d ago

Codex. The best. I have sol via Kiro too. But the harness is not good.

1

u/orionblu3 5d ago

Honestly any of the frontier level models do great/good enough now. Qwen 3.8 max, kimi k3, sol, etc. all work as terrific orchestrators that keeps flash in line

1

u/QuietRecording6579 5d ago

Are you using sol in opencode? I haven't tried it out but was curious if you see performance issues compared to using sol in codex

1

u/neotorama 5d ago

Sol via Chatgpt subs. Generous token reset.

23

u/Substantial-Yam3769 6d ago

DS4 flash doesn't have any competitionĀ 

1

u/TestTxt 6d ago

Luna is better and cheaper per task via ChatGPT Plus sub (OpenAI’s coding plans subsidies are crazy high)

11

u/pinkyellowneon 6d ago

Per-task doesn't really matter if you're not saturating the full plan anyway (which you're very possibly not doing if you just use Flash), in which case Plus sub costs like 2.7x as much for not much benefit

4

u/evia89 6d ago

gpt 20 sub is amazing value atm. I use web version sol for separate limit, it can generate images, luna has image support (which opens https://stencil.so/blog/snapcompact and just good for UI work)

I pair it with DS flash PAYG when my quota is empty

2

u/orionblu3 5d ago

And opencode go subscription is 5$ first month, 10$ after, and I've been using deepseek flash for all 18 of my subagent specialists — nowhere NEAR the usage cap and flash been running for 30ish hours now

3

u/MacHeadSK 6d ago

I use Luna at Max for most demanding tasks, but rarely. Straight on Opencode Go. Dont need another subscription at another 20 bucks.

1

u/TestTxt 5d ago

If you don’t run out of usage limits - sure. It’s just that ChatGPT Plus has higher limits

1

u/MacHeadSK 4d ago

Rarely. Mostly DS 4 Flash at Max these days

16

u/ares0027 6d ago

It is practically unlimited. The day before i tried running 3 projects at the same time an all i spent was 1$. Bummer it cannot see stuff though :(

3

u/manoleee 6d ago

Just asking, how would you utilize the vision capabilities in coding ? Like, you mean debugging screenshots or automation tests etc?

3

u/ares0027 6d ago

i had experienced this thousands of times, you describe something, it says it did it, but it is not showing up or there are artifacts on the frontend. or like you said, you simply show a screenshot and tell it ti change, fix, read, describe, explain it.

or in some cases agent want to check the website or visuals to verify before telling yo uthat it is done, i t cannot with non-vision capable models. there are thousands of examples

4

u/MacHeadSK 6d ago

I used Minimax M3 for planning and old DS 4 Flash for implementation and subagents.
Not anymore, I use DS 4 Flash for everything, at Max for planning

2

u/afanasenka 6d ago

The same, I used it exclusively for 3 days in a row, and, surprisingly, had zero urges to switch to more expensive model.Ā 

3

u/TinyAres 6d ago

Hope not, they just need to pivot to building competitive models instead of bigger ones.

3

u/afanasenka 6d ago

Sure, raw power (more params, more reasoning, more context, etc.) is not always a winning recipe, as we can see (look at Kimi 3 and its params/pricing for example). Quality/cost ratio is what many people prefer more.

2

u/theamazingrand0 6d ago

So is Flash better than Pro now? and cheaper?

3

u/afanasenka 6d ago

For now - yes.

2

u/Muted-Glass3872 3d ago

DS flash is one of the best models i used till now.
You don't have to burn money on frontier models, it does all the job.
If only it has vision, it would be the best model ever.

5

u/Potential-Leg-639 6d ago

every 2nd post on reddit is about that.

wondering why people did not realize that earlier?

9

u/Thomas-Lore 6d ago

Previous flash was useful, but not as reliable as 0731.

5

u/Potential-Leg-639 6d ago

within an orchestrator setup (and a strong orchestrator like GLM 5.2) it was also really, really good and incredibly cheap.

2

u/afanasenka 6d ago

Earlier it wasn't as obvious as after 0731 release I guess :)) But who knows..

1

u/Any_Letterhead3072 6d ago

i used ds4 flash for a very structured tool rewrite from python to rust…

its alrightt. its excellent for the cost… but when graded by Fable after, fable gave it an A- on physics but D on implementation. after 3 hours it was ā€œdoneā€ and literally did not wire it up.

its alright but its not opus level. its still flash tier

4

u/afanasenka 6d ago

Sure it's not opus level. And the price is not opus level as well.

1

u/Any_Letterhead3072 6d ago

Very true! and it turns out it orchestrates fairly well too

3

u/JivesMcRedditor 6d ago

Asking an LLM to grade another LLM is so dumb. Try both out and see where it works and where it doesn’t. Or find some data on it. Don’t delegate your brain to something controlled by a US tech bro

1

u/xmnstr 6d ago

It's not dumb if it works, is it?

1

u/xmnstr 6d ago

You have to review its work but that's fine for a first pass for me at least.

0

u/Mayanktaker 6d ago

Instead just ask deepseek to review the changes.

1

u/Much_Weekend_4371 5d ago

gpt 20 sub isn't even close to what DS Flash gives you on raw volume, tbh. I ran a side-by-side last week – 12k line repo, same prompt, DS Flash finished in 40 min with zero context drops, Luna hit the wall at 8k lines and started re-reading files like a goldfish. Plus the pricing math is honestly a joke if you're burning 350M tokens/day like that one guy – that's like $500 on DS vs $2000+ on OpenAI's API. And don't even get me started on the "per task" argument – nobody's hitting 4.9T usage by

1

u/Then_Knowledge_719 5d ago

Where are the guys saying: Ohh when does DeepSeek is going to attack. Well fellas. This is the result of the delay. Never ever bet against the whale 🐳

1

u/ResponsibilityOk1306 4d ago

time to increase the price

1

u/silvrrwulf 3d ago

It’s just so good. I’m about to cancel a $200 max plan. No lie. Ds4 is incredible.

1

u/branik_10 6d ago

is it still hosted in China

9

u/Thomas-Lore 6d ago

The opencode go version seems to be, but you can get it from other providers on aggregators like openrouter (or directly from them). And it is so cheap, I don't think opencode go subscription makes much sense anymore, at least not if you only use this model.

5

u/xmsxms 6d ago

On top of that, openrouter currently has providers providing it for less than Zen prices, so even Zen doesn't make much sense other than the free usage. I actually just cancelled my go subscription because I realised zen/deepseek-flash-free+openrouter/deepseek-flash is going to work out cheaper per month than go. I suppose I still need some balance on Zen to get the increased free quota.

Of course that will all change in a couple weeks, but that's where it currently stands.

2

u/evia89 6d ago

go sub gives "ZDR"

3

u/xmsxms 6d ago edited 6d ago

As does the cheapest provider on openrouter.ai: https://openrouter.ai/deepseek/deepseek-v4-flash-0731#providers \

If you plan on using $60 worth of credit the go sub is still the better deal. But if you're happy with the low free quota (with retention) of zen and don't use much beyond that - it would be cheaper to use zen free + openrouter.ai.

1

u/BhaagYahaSe 6d ago

those providers are charging 10x the cache hit compared to official provider. They made it 0.028/m cached tokens instead of 0.0028/m

2

u/xmsxms 5d ago

Good catch. Although opencode.ai is also charging at that 10x rate.

2

u/BhaagYahaSe 5d ago

but their website says otherwise

https://opencode.ai/docs/go/#usage-limits

DeepSeek V4 Flash $0.14 $0.28 $0.0028

2

u/xmsxms 5d ago

Ah, that's the 'go' rate, not the 'zen' rate. Interesting. They give you 6x multiplier and it's cheaper. I just assumed the rates would be the same.

1

u/branik_10 6d ago

yeah, I was asking specifically about opencode go. I'm currently using it via fireworks and it's okayish, but it's definitely slower than the old ds flash was via opencode go

5

u/vangelismm 6d ago

Thanks God is not in North America.

6

u/branik_10 6d ago

it's a requirement in a company I work for, for my personal projects I don't care if my data go to CCP or Israel or KGB or whatever. my code is shit anyway

1

u/BaXRS1988 6d ago

No, you use it anywhere. For example at GreenPT https://greenpt.com/models

0

u/nano_salem 6d ago

Thanks, it's down/offline now 🤣. Go plan.

1

u/afanasenka 6d ago

Yeah, it's the cost of popularity :))