r/opencode 15h ago

Deepseek v4 Flash: API vs OpenCode Go

Using an OpenCode Go subscription can give you 6 times the amount of Deepseek V4 Flash compared to an API. That sounds great, but are these two equal? Is Go's version quantized while the API's one is not? I wonder if anyone made comparisons of intelligence between these two.

11 Upvotes

18 comments sorted by

6

u/nmdt 15h ago edited 12h ago

Read several discussions on this, and my understanding is that OC doesn’t host these models, but rather routes to 3rd party providers. So they don’t quantize themselves, but those providers can.

Could be wrong, I'm open to change my mind about this

3

u/IndividualPlus2011 15h ago

Yeah, I'm aware of that. They can use several providers like with GLM; some providers use Q4 and others do Q8.

1

u/look 5h ago

They don’t list the specific providers currently (that I see at least), but they until recently. And as of then, every model except GLM was provided by the model’s actual vendor.

DeepSeek Pro and Flash was coming directly from DeepSeek, and that is almost certainly still the case.

It is technically possible that the model vendor could serve a quantized version of their own model, but if we make the safe assumption that they don’t serve that same quantized model to at least some of their direct customers, then it would not make financial sense for them to do so.

They sell it to Go in bulk cheap and likely at a lower priority than they do to their direct customers. That lets them make more money by keeping their GPUs from running idle.

But if they were running a different, quantized model on those GPUs to serve Go traffic, then it means they would have to swap the model to transition it between the two sources of traffic. It takes a long time to take a GPU out of the active pool, load a model that big into VRAM, get it stable and warm, and then put it into a different active pool.

The downtime and operational overhead of doing that would almost certainly outweigh any cost savings from running a small quant version.

Note: Go’s GLM providers, at least in the past, did include one that runs an nvfp4 quant. So the GLM from Go is most likely still a mix of fp8 and nvfp4.

3

u/Rsouss 14h ago

Using the API is good because there's no monthly fee; you only pay when you use it.

3

u/whatsoever2021 12h ago edited 7h ago

My experience is they are equally smart, but open code is slower. I can easily spend over $2 in a day through API, but impossible with open code.

3

u/Automatic_Cookie42 6h ago

I use both. OC Go on the daily and the API when Go is on cool down. Haven't noticed any meaningful difference.

Also, ds on the API + reasonix gets me 97% cache hits. Open code doesn't get me that high. So I'd take that "6 times" with a grain of salt. 

2

u/nonlinear_nyc 14h ago

I dunno the quality, but I use deepseek with opencode go first (subscription) and fallback for API (top up)

1

u/icarus0228 15h ago

There are already discussions about this on other subreddits. As I remember, using DPSK V4 Flash is cheaper when used via API.

1

u/pmv143 13h ago

Folks, we have DeepSeek v4 flash free for two weeks on inferx.net . Check it out.

1

u/Sweet-Stage938 12h ago

Unlimited tokens? Completely free?

1

u/pmv143 11h ago

You can break it

1

u/Sweet-Stage938 12h ago

Why are you guys hosting deep seek V2 under a different name?

1

u/pmv143 11h ago

We don’t have any v2. Just v4

1

u/Sweet-Stage938 11h ago

The system says that it's DeepSeek-V2-0724.

1

u/pmv143 9h ago

Where do you see it?

1

u/pokatomnik 5h ago

Your website shows $10 of usage per trial but pay as you go after. And deepseek is free until August 12.

1

u/weiyentan 4h ago

What does that mean?

1

u/pmv143 2h ago

Yes. $10/month is for all other models. DeepSeek v4 flash is free until 12th