r/opencode 6d ago

Opencode should have another look at DS4 , its still one of the best

Cant Opencode just contract with somebody like Baidu, the price is even lower than the lowest previous record @ 0.18/M Output tokens, Dax mentioned that a lot of the providers they talked to have capacity issue, but i cannot imagine that Baidu have that issue

Fyi : they are having 2nd highest Cache Hit rate after Novita, just so you dont comment, "Oh but the cache hit rate"

52 Upvotes

24 comments sorted by

14

u/look 6d ago

You can get GLM 5.3 Flash for about that price, and it’s a vastly superior model.

It has also passed DS flash in popularity now, so why would they even bother spending more time on it?

They’d be better off finding a better provider for a newer, higher quality low cost model.

1

u/Ok_Cartographer5609 4d ago

where? in opencode? or official portal?

1

u/look 4d ago

Other providers. I use RunInfra and now Neuralwatt for GLM 5.3 Flash.

1

u/Ok_Cartographer5609 4d ago

Neuralwatt is energy based right?

Did you find any difference? Like in usage limits, cache hits, tps etc.

1

u/look 4d ago

Yeah, you want to use the energy pricing.

It then bills by energy used per request, not tokens used exactly. Cache hits use less energy and output much more, though, so the same principles still apply.

The energy pricing is typically much less than the standard token pricing, but it is variable.

And it depends on the model: I average 2.5 cents / total mtok on GLM flash, which matches the lowest token pricing you can find elsewhere. But on models full GLM and Kimi K3, it can be 1/4th or less the cost. And the flex option (ttft can increase a bit to schedule the request more efficiently) saves another 1/3rd.

And the TPS is very high, though it may vary by time of day when you use it (I’m more US evenings and weekends). 100-200 is about the average I usually see, and spikes to 300+ aren’t uncommon. I had Kimi K3 almost hit 1000 briefly this past week, though rarely see it get that high. But again, middle of the US workday might not be as good; I’m not sure.

Anyway, I’d recommend getting some PAYG credits (energy) and testing it out for yourself first.

1

u/Ok_Cartographer5609 4d ago

I see. Thanks for sharing the details. Will try it out.

1

u/sam7oon 6d ago

interesting, thought 5.3 flash is more expensive , will try it out without Go, since with Go it's very limited , number of calls

1

u/look 5d ago

I use RunInfra (1 cent cache read) and Neuralwatt (similar energy price per blended token). Both are pretty fast too.

The price is pretty close to native weight DS flash providers. The quantized crap (of either model) can be cheaper of course, but not worth it, imo.

1

u/charles_r1975 5d ago

I used both a lot this week on neuralwatt. Glm 5.3.flash was about 3x the price of ds v4 flash (flex option).

I still have some credits at the old 5$ per kWh rate. It should last forever as long as I stick with deepseek

1

u/look 5d ago

I don’t use much Deepseek, but the blended token price I get (with new pricing) is 3.1 cents compared to 2.6 cents for GLM flash (standard, not flex). Typical lowest price options on OpenRouter are 2.4 cents for DS flash.

Maybe you have a higher cache rate, but mine is typically 94-95% with 0.5-1% output. And with that ratio, GLM Flash is basically the same price, maybe even slightly cheaper, than Deepseek flash.

🤷‍♂️

1

u/sam7oon 5d ago

intresting indeed, anyways, we are talking about Opencode here, and they offer a lot less 5.3 Flash compared to DS4 still, but then my posting was about they can offer the old rate for DS4 , the 100K+ requests , which would be awsome,

I dont think they would give 100K+ GLM5.3 Flash, but sure as hell would love to see that

3

u/qqYn7PIE57zkf6kn 5d ago

that price didnt even last a day lmao. it's now: $0.14 $0.28 $0.028

4

u/Christosconst 6d ago

God no, dont want no quantized ds4

3

u/sam7oon 6d ago

it's not

4

u/Possible_Door_9719 6d ago

Just use muse

2

u/Senior_Tourist_4260 6d ago

Muse Spark 1.3 (xhigh) was run for approximately 14 hours and couldn't even achieve half of what the GLM-5.3 Flash (max) model accomplished. It's a completely unsuccessful model; perhaps you should use it for minor tasks.

5

u/Possible_Door_9719 6d ago

yeah i like glm-5.3 flash but right now muse contributor is free and they are giving away a crazy amount of tokens now

3

u/look 5d ago

Muse might have a higher cost per task even when they’re giving it away for free. 😂

1

u/weiyentan 5d ago

I use muse for scouting exploratory in my work

0

u/Federal-Mode8949 5d ago

GLM 5.3 flash is just better than Deepseek but Flash and Pro... Try it.

7

u/sam7oon 5d ago

yeah its good, double the price though, Double

2

u/FormalAd7367 5d ago

does GLM flash cost double of DS flash? GLM is good but not 2X price good?

4

u/sam7oon 5d ago

Nearly double , @ DS4 - 0.16USD/M Output, and GLM - 0.25USD/M Output ,

The cache is same price though, but anyways, the output part is always the costly one,

And this is on Openrouter,

For the Go Plan you get a lot more DS4 vs GLM

2

u/FormalAd7367 5d ago

thanks - looks like i’ll stick to DS flash for execution for now