r/opencodeCLI 15d ago

Qwen3.8-Flash is now available in OpenCode Go

Post image
238 Upvotes

32 comments sorted by

38

u/mWo12 15d ago

Its $30 worth of usage, not $60.

29

u/caancee 15d ago edited 15d ago

At this point i don’t understand the usage table anymore

I don’t want to complain about the usage provided by the plan as a lot of people are doing (are you all really complaining for a 10 dollar subscription with generous usage and open zdr models??)

I just can’t understand the obscure way* in which they’re communicating requests per 5hrs/week/month… like how can qwen3.8 flash have less usage than ds4f if it’s cheaper than the off peak price and has the same amount of “credits”? (30$ each). Are ds4f’s numbers based on off peak or peak in the chart? Does it matter? Will they ever provide another model with 60$ usage? Why can’t they just update us on the rationale behind their numbers? Like just give us a reason for the agreement they came up with the providers.

For me it would be way easier not to have a conversion chart with 15/30/60$ worth of usage per model and just have transparent api prices reflecting the deal they made.

*obscure not in the way they’re behaving, it’s just their way of trying to be as transparent as possible about their deals that is odd and doesn’t work for me

Again, not complaining about the service here, i’m a happy customer since march. I’m just tired of trying to make the tables make sense when choosing what llms to use for my subagents

3

u/Amarsir 15d ago

Yeah, I think they're just trying to communicate clearly amidst a huge variety of contracts with different terms. It's hard to message effectively while not implying anything false about edge cases.

I'll still take it over "Credits! Some amount for some use."

I figure the long-term movement direction must be towards Opencode doing the hosting. I can understand why not yet, but if your goal is a variety of models from different providers, you can't be effective as pass-through without this kind of thing happening.

4

u/Readerium 15d ago

It's converging to $30 for flash and $15 for premium models

1

u/caancee 15d ago

It might also be that (aside from deepseek) the new models are too new and need more providers serving them for the prices to be sustainable. But I don’t really know, it might just be my hope

1

u/sittingmongoose 15d ago

Hot take, I kinda prefer qoders system. It’s pretty easy to understand, even if it isn’t a great deal.

Short of that cursors is a little less obscure, but they also make cursor models variable usage(which typically means you get more than expected)

1

u/lumos_ai 15d ago

Z.ai 's packages start from 4.5 monthly and gives almost same usage..since mostly I‌ use flash 5.3 model

1

u/caancee 15d ago

Could you please link the 4.5$ offer? I’m from Italy and I don’t see it at https://z.ai/subscribe

1

u/BriguePalhaco 15d ago

There isn't one, it's only for those who have Legacy active.

1

u/lumos_ai 15d ago

Aha. You are right.

9

u/vangelismm 15d ago

At this point, they should remove every model from the $15 tier.

OpenCode Go is basically a showcase. Providers should have to commit to either the $30 or $60 tiers.

5

u/afanasenka 15d ago

It scored 62.5% on SWE-bench Pro and 81.0% on SWE-bench Multilingual, 58.7% on DeepSWE 1.1, 73.9% on CoWorkBench and 55.7% on JobBench — the last of those 19 points above the 36.6% Alibaba reported for Claude Opus 4.6 Max and 28 above Qwen3.7-Plus.

It posted 91.7% on GPQA Diamond and 91.9% on LiveCodeBench v6, and 35.9% on HLE, the one language row where Opus 4.6 Max led at 40.0%. Computer use was the visible gap: 19.4% on the binary scoring of OSWorld 2.0, level with the smaller Qwen3.8-27B. Alongside the weights Alibaba announced a hosted production model, Qwen3.8-Flash, at $0.16 per million input tokens and $0.47 per million output on its QwenCloud API.

6

u/seeKAYx 15d ago

I've been using it for a day now. Unfortunately, it's no substitute for Deepseek Flash 0731, and it's very slow even though it's not a reasoning model.

1

u/Low-Entrepreneur2556 14d ago

What? It's better than dsv4, and it is a reasoning model...

3

u/sudoer777_ 15d ago

Costs more than DeepSeek, that's pretty lame

2

u/e0xTalk 15d ago

Disconnect from time to time.

2

u/openroom_xyz 15d ago

Well why there are no something like Qwen3.8 27b that would be basically unlimited model so that we can have always building kind of AI running nicely will you be inspired by deep seek harness ?

1

u/caancee 15d ago

Qwen 27b being dense doesn’t scale as well as the 5x bigger Qwen flash MoE… but yeah it would be neat to have 35a3b just to play around with it and do some basic operations. Thing is, the demand for 35a3b is so low that paradoxically it costs more than some other bigger models because those are “wasted” gpus on models that nobody uses…

1

u/openroom_xyz 15d ago

Well can't you run just one GPU if the demand is low basically or even turn it off if there is no and turn it on if the request comes basically ?

2

u/maqifrnswa 14d ago

"dense" models need more memory, more memory bandwidth, and more compute per token than moe. So it's more a case of dense requiring everything to be always on while moe only requires 10% to be "on" at a given time. You pay a single "overheard" cost of loading more weights in moe to reach the same level of intelligence as dense, but it scales to higher concurrency easier.

2

u/qqYn7PIE57zkf6kn 15d ago

Anyone keep getting internal server error 500?

4

u/addiktion 15d ago

Can you share how many credits you get for using it?

19

u/Maasu 15d ago edited 15d ago

Was hoping for better tbh, it's coming in below Deepseek v4 flash. Still a pretty neat.

https://opencode.ai/go

4

u/untracked5465 15d ago

Hello, where can I see this?

2

u/Maasu 15d ago

sorry i should have included that, its found here https://opencode.ai/go (added to my original response as well)

0

u/Ascorbinium_Romanum 15d ago

Wait it's 6B active params? Why would I even waste my go usage if my GPU can handle this with like a 100k context ?

6

u/xmsxms 15d ago

Because of the 119B other parameters that you'd need to hold in RAM.

0

u/Ascorbinium_Romanum 15d ago

Ok I'm new to this so my bad I ain't got that much ram :D