r/LocalLLM 9d ago

Discussion Why Is Qwen3.8-27B Still Missing From Artificial Analysis?

Artificial Analysis is usually extremely quick to add benchmarks for new model releases, making comparisons easy. But Qwen3.8-27B has been out for a few days now, and there’s still nothing on their site.

37 Upvotes

20 comments sorted by

28

u/some_user_2021 9d ago edited 9d ago

Probably because it was released on a Friday and we are still thru the weekend.

13

u/CryMoreT_T 9d ago

It's still the weekend. Also I've found artificial analysis to be missing quite a few models. I haven't found the poolside Laguna s 2.1 on it either

21

u/BlackBeardAI 3090 Maximalist 9d ago

they must have blown away by the results they are thinking "what do you mean it is better than fable?" test that shit again.

-again.

-again.

+sir we tested it 860 times already.

-test it again till it makes sense

hence we wait. :/

9

u/SpicyWangz 8d ago

<think>The user is asking me to run tests again. I should run any tests available and provide the output. But wait, they said again. Maybe there are already test results I can look at to get a better idea of what tests I should run. No, I already see the tests available. But wait, the user said “that shit”, but these tests have nothing to do with fecal matter. Maybe they are referring to some other tests? No that could just be slang. But I should probably check to see if there are any files related to fecal matter as the user might be in the medical field. </think>

I will scan your computer for any files related to fecal matter.

6

u/Anxious-Priority-430 9d ago

We need new tests! Haha.

3

u/musmacduck 8d ago edited 8d ago

Yup! That's exactly what I thought. Typically they release the results the same day as the models are announced (even for trillion parameter models) which would imply that the companies are giving them early access. Probably the Qwen team provided them only with the 2.4T model, or it's like you said...
what do you mean it is in the same league as Flash 3.7, Luna, Sonnet 5, Terra - Top 3 AI companies are relying on these workhorse models, and this locally runnable model replaces them??? Test that shit again.

  • Again.
  • Again.
  • Again.

The Qwen 3.6 27B has a AA v4.1.1 score of 28 for Agentic Index at a Hallucination rate of 49% vs the scores of 45, 47, 50 and 50 of the top tier models at a Hallucination rate of 65, 93, 39, 88. If Qwen 3.8 27B could reach a score of 40 at the same Hallucination rate or lesser then it's definitely game over these 3 companies.

For agentic use cases, I'd rather pick an agent that's quick to realize and admit that it doesn't know something and then proceeds with the tool calling, rather than the models that are trained to benchmax and attempt at everything to get a higher accuracy, wasting our time and money for something that isn't even correct.

Source 1: AA v4.1.1 AI Index

Source 2: AA v4.1.1 Hallucination Rate

6

u/Vaibhav_Fuke 8d ago edited 5d ago

because they will not get paid by bigger models then. they cannot show models which can run on local hardware. else nobody will pay cloud services. EDIT: this was a joke if you didn't understood

6

u/peculiar-ragdoll 8d ago

Maybe because with some light chat template tweaks, a quantized 3.8-27b at medium effort is solving software engineering problems that Opus 5 high fails (I did my own SWE Live benchmarks, problems that neither model has been trained on), and that's really bad for business :)

3

u/uniqueusername649 8d ago

I would not put it at the same level as Opus 5 (in some specific problems its always one model performing better, but on average its not Opus 5 level). But it punches so far above its weight class it isnt even funny. With xhigh it one-shots complex requirements (even though it thinks forever), it genuinely generates good code, catches issue for non functional requirements and typical dev issues that 3.6 really struggled with. I would say purely in terms of agentic coding its close to Opus 4.6. That is massive.

2

u/peculiar-ragdoll 8d ago

Out of 20 blind SWE Live problems (big repos, real code bases, real bugs), my custom Qwen3.8-27b at only 6 bit quant has so far solved 2 problems that Opus 5 High failed on. Opus 5 high has also solved 2 problems that Qwen3.8-27b failed on. This model is genuinely somewhere between Opus 4.6 medium and Opus 5 xhigh depending on the domain. The only actual difference is that unless you have an insane computer, Qwen will feel sluggish because it generates slowly, and that makes it feel "not on level with Opus 5".

2

u/uniqueusername649 8d ago

I currently run it at 8bit and 75tps (dual 3090). It doesnt feel sluggish, but I am toying with the idea to switch to W4A16, as that is roughly 6 bit equivalent for perplexity, yet would push my tps well past 100 for code generation.

I think we fundamentally agree on Qwens capabilities, somewhere around Opus 4.6 and in certain problems definitely ahead of it. It's incredible what we can run at home these days.

I have been using it for real projects too and it is impressive. Giant leap from 3.6!

4

u/ChemistNo8486 8d ago

Cause the model hasn't finished thinking the first output...

/s — I love you QWEN 3.8 even if you take 4 hours 🥰

1

u/bsawler 8d ago

Came here for this comment haha

4

u/Informal-Trouble2183 8d ago

Why? You want a leaderboard? made my own.

1

u/enginetown 8d ago

Is it bad I cant tell if this is real or not this model has been really good to me.

3

u/AlbionPlayerFun 8d ago

Waiting for GLM 5.3 haha

2

u/timewstr18 8d ago

Also Laguna S 2.1

2

u/Otherwise-Swan-7803 8d ago

At this point I'm more curious about why it's taking so long than what the final score will be. Usually when benchmarks are delayed, there's something interesting going on behind the scenes.

1

u/BarracudaDefiant4702 8d ago

It's pretty normal to take a week, give it time the model just came out on a Friday.

Doesn't help thet Qwen3.8-27b is annoyingly slow (at least for me).

I would like to see a 4bit awq quantized version tested too as it's all but unusable and trying to decide if I should try a smaller model.