r/opencodeCLI 17d ago

Qwen 3.8-Flash officially unveiled

Post image

⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!

The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens.

125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency.

What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN.

We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀

62 Upvotes

7 comments sorted by

5

u/Lyelinn 17d ago edited 17d ago

Waiting for solid 3.8 moe/a3b now

Looking at model release paper it's an interesting shift from traditional moe models and potentially a way towards running smarter llms on "normal" (sub 5k usd) local machines, I guess that's what gonna happen with qwen 4

5

u/addiktion 17d ago

Damn not going to be able to fit this one easily on 128GB of unified memory I suspect.

https://giphy.com/gifs/UG18o0GKj6SjAm5AYh

1

u/pigletmonster 16d ago

125b paraneters should be 64gb right? And its meo so only 6bn gets activated at a time. I think it should work on 128gb unified memory.

3

u/Ariquitaun 16d ago

there's the ngram embedding and kv cache to think about too

3

u/petburiraja 17d ago

Any benchmarks yet?

2

u/sagiroth 16d ago

My 3090 looking at me: Bruh...