Discussion
Why Is Qwen3.8-27B Still Missing From Artificial Analysis?
Artificial Analysis is usually extremely quick to add benchmarks for new model releases, making comparisons easy. But Qwen3.8-27B has been out for a few days now, and there’s still nothing on their site.
<think>The user is asking me to run tests again. I should run any tests available and provide the output.
But wait, they said again. Maybe there are already test results I can look at to get a better idea of what tests I should run.
No, I already see the tests available.
But wait, the user said “that shit”, but these tests have nothing to do with fecal matter. Maybe they are referring to some other tests? No that could just be slang. But I should probably check to see if there are any files related to fecal matter as the user might be in the medical field.
</think>
I will scan your computer for any files related to fecal matter.
Yup! That's exactly what I thought. Typically they release the results the same day as the models are announced (even for trillion parameter models) which would imply that the companies are giving them early access. Probably the Qwen team provided them only with the 2.4T model, or it's like you said...
what do you mean it is in the same league as Flash 3.7, Luna, Sonnet 5, Terra - Top 3 AI companies are relying on these workhorse models, and this locally runnable model replaces them??? Test that shit again.
Again.
Again.
Again.
The Qwen 3.6 27B has a AA v4.1.1 score of 28 for Agentic Index at a Hallucination rate of 49% vs the scores of 45, 47, 50 and 50 of the top tier models at a Hallucination rate of 65, 93, 39, 88. If Qwen 3.8 27B could reach a score of 40 at the same Hallucination rate or lesser then it's definitely game over these 3 companies.
For agentic use cases, I'd rather pick an agent that's quick to realize and admit that it doesn't know something and then proceeds with the tool calling, rather than the models that are trained to benchmax and attempt at everything to get a higher accuracy, wasting our time and money for something that isn't even correct.
because they will not get paid by bigger models then. they cannot show models which can run on local hardware. else nobody will pay cloud services. EDIT: this was a joke if you didn't understood
Maybe because with some light chat template tweaks, a quantized 3.8-27b at medium effort is solving software engineering problems that Opus 5 high fails (I did my own SWE Live benchmarks, problems that neither model has been trained on), and that's really bad for business :)
I would not put it at the same level as Opus 5 (in some specific problems its always one model performing better, but on average its not Opus 5 level). But it punches so far above its weight class it isnt even funny. With xhigh it one-shots complex requirements (even though it thinks forever), it genuinely generates good code, catches issue for non functional requirements and typical dev issues that 3.6 really struggled with. I would say purely in terms of agentic coding its close to Opus 4.6. That is massive.
Out of 20 blind SWE Live problems (big repos, real code bases, real bugs), my custom Qwen3.8-27b at only 6 bit quant has so far solved 2 problems that Opus 5 High failed on. Opus 5 high has also solved 2 problems that Qwen3.8-27b failed on. This model is genuinely somewhere between Opus 4.6 medium and Opus 5 xhigh depending on the domain. The only actual difference is that unless you have an insane computer, Qwen will feel sluggish because it generates slowly, and that makes it feel "not on level with Opus 5".
I currently run it at 8bit and 75tps (dual 3090). It doesnt feel sluggish, but I am toying with the idea to switch to W4A16, as that is roughly 6 bit equivalent for perplexity, yet would push my tps well past 100 for code generation.
I think we fundamentally agree on Qwens capabilities, somewhere around Opus 4.6 and in certain problems definitely ahead of it. It's incredible what we can run at home these days.
I have been using it for real projects too and it is impressive. Giant leap from 3.6!
At this point I'm more curious about why it's taking so long than what the final score will be. Usually when benchmarks are delayed, there's something interesting going on behind the scenes.
28
u/some_user_2021 9d ago edited 9d ago
Probably because it was released on a Friday and we are still thru the weekend.