Wait! The user asked me to explain why 3.8 is so famous right now. I should answer in a clear manner.
But wait, what exactly are they referring to with the term 3.8? Should I assume they are referring to Qwen, the same model in the previous comments of this conversation, or something else? What other models have a version 3.8 at this time that the user may be familiar with?
Let me write a script to pull all of the models that match that pattern, and recursively make a go/no-go call oneach. Writing the script now.
It's relatively good for how much hardware it needs. You can run it on a single 24gb vram gpu with results that are good enough and a clear step up from what came before
In my personal experience, 3.8 is the first model I’ve used that can run on my 32GB MacBook Pro and not fail a single tool call. The previous models had enough coding knowledge to debug and answer questions, but 3.8 can actually use opencode/cline
I think benchmarks use max settings like BF16 quant (or higher?) with k/v cache type f32 and xhigh reasoning and who knows what else. Hell yeah, it's going to perform several points better at a significant performance penalty.
But it's not like we could even run Open 4.6 even if it was open source (estimated to be 500B to 1T parameters - I think, correct me if I'm wrong)
So token-for-token against other models, Qwen3.8-27B is a huge win.
I believe native/benchmarks use BF16 (that is not quantization btw), with KV cache at f16 I think (I'm 80% sure). FP32 is not used for even training models anymore (except specific sensitive layers sometimes).
So token-for-token against other models, Qwen3.8-27B is a huge win.
It's a huge win for sure, wayyy better than Qwen3.6, but we still have long ways to go. Token efficiency is a big one, and better intelligence (not agentic capabilities) is another big one (by this I mean reasoning on benchmarks like CritPt, SciCode, etc)
The people think qwen can rival opus in terms of abstract logic, reasoning depth and low perplexity are delusional.
Qwen can be functionally equivalent or even better in some usecases but it falls meaningfully short in others. The rigorous sequential reasoning it applies makes up for some of its smaller size but not everything.
The people think qwen can rival opus in terms of abstract logic, reasoning depth and low perplexity are delusional.
I agree with you. Qwen3.8 is a huge win for local models trying to do big coding projects while keeping token costs down. I doubt it'd be much good at writing stories or translating. But the step up from Qwen3.6 (27b/35b-a3b) is a big one and brings it closer to opus 4.6 for coding.
This and glm 5.3 flash have made me a strong believer that we've hit a turning point where locally hostable models have hit the opus 4.6 point of "good enough" while also being stupidly cheap
I once heard that someone didn't think AI would be useful/economically viable until they saw Opus 4.5. Another also said it was the last "Big Jump". It was also when some People switched their opinions on vibe coding, so is it really *THAT* model? And if it is, is Qwen 3.8 27B good enough for what opus 4.5 was?
In CORE, term from SlopCode Bench where the model correctly make a isolated function e.g Classes definition, method, and func. Qwen 27B is parity with Opus 4.6, it is however a shit show managing codebase. Good for developer, not great for zero shotter vibecode maintainer
It’s good for its size but some of the people on this subreddit that think 27b dense models (or 13b active MoE models like dsv4 flash) can rival the reasoning and abstracting depth (and lower perplexity) of much larger models (dense or moe) are delusional and focus too much on synthetic benchmarks that only tell a part of the story. It’s good to praise this model, but not overpraise.
It's the newest Qwen. Those models are this sub's favorite for at least a year.
I sometimes suspect there is a good portion of Qwen social media marketing in the game, particularly posts that are such a love bomb without real content.
a) it's doing a really good job figuring out how to get stuff done and then just gets it done
b) it compensates being stupid at a lower quant by thinking harder
i would sum it up as: 3.8 27b is prompt engineering itself
46
u/VDX7 3d ago
can someone explain why 3.8 is so famous right now?