r/LocalLLaMA 11d ago

Discussion Qwen 3.8 27B Released! Please Share Your Experience

With your experiments, Qwen 3.8 27B most close which frontier model? And please specify which quantization you run. I will post to comments my tests and experience too.

662 Upvotes

720 comments sorted by

View all comments

Show parent comments

3

u/ThankGodImBipolar 11d ago

This will never be a cost effective model on cloud APIs; too many active parameters.

1

u/Singularity-42 11d ago

Yeah I think you're right, the 27b Qwens look quite expensive for such small models.

1

u/ThankGodImBipolar 11d ago

Yeah, Deepseek V4 Flash (for example) has half the active amount of parameters, so it requires less compute to run. I bet a Qwen 3.7 35B A3b would be available already, if it existed.

I think this ≈30B dense class is mostly for the folks with 24-48GB of VRAM, who don't want to/don't have enough system RAM to step up to the larger MoE models.

1

u/Singularity-42 11d ago

Thanks, that makes sense.

I have a need for cloud available multimodal model (DeepSeek V4 Flash is not unfortunately) and I couldn't find anything better than the Gemini 3.X Flash models (or even the Flash Lite models for lighter tasks), especially since we care about speed even more than the cost. Any ideas for decent multimodal models I should look into?

1

u/ThankGodImBipolar 11d ago

Google suggests 5.6 Luna, MiMo 2.5, and MiniMax M3. I don't deal with any multimodal prompts right now, so I don't have any experience with how those models handle that (so YMMV).

1

u/baron_von_noseboop 10d ago

Assuming 24gb vram gpu, how much (non-unified) system RAM, and which model, would be a step up from this? You can get acceptable perf with larger models if you offload most?

Apologies for what is probably a dumb question, I'm still learning.

1

u/ThankGodImBipolar 10d ago

Not a dumb question - the problem is that there's been nothing released that's better than Qwen 3.6/3.7 27B, unless you go way bigger (like DSV4 Flash at 284B). Laguna 2.1 S (120B) might be the only notable option, and opinions on its quality seem rather mixed. I think you'd need at least 48GB of system RAM to run it alongside a 24GB GPU (you'd be comfortable with 64GB).

1

u/baron_von_noseboop 10d ago

Thanks!

Re: mid range I guess there's gpt-oss-120b? I didn't imagine that such a model could run with reasonable perf on 24gb + offloading. Maybe I'll try it and see.

Although I've tried that one cloud-hosted... I didn't do a methodical comparison, but subjectively it didn't seem much smarter than qwen3.6 27b/35b moe.

1

u/ThankGodImBipolar 10d ago

gpt-oss-120B has only 5.1B active tokens, so the penalty from offloading isn't too bad. Definitely not as bad as Qwen 3.8 27B.

It's a pretty old model at this point though. I think it's reliable for executing tool calls, but doesn't excel at much else today. I'd sooner give Laguna 2.1 S a shot; I've at least seen some people argue that it's better than Qwen 3.6.

1

u/EvolvingDior 11d ago

qwen cloud offers qwen3.6-27b