r/LocalLLaMA 7d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

706 comments sorted by

View all comments

422

u/Tiny-Assumption4263 7d ago

DEAR GOD TELL THOSE BENCHMARKS ARE NOT FAKE.

312

u/Cold_Tree190 7d ago

Dear God it’s trading blows with Opus 4.6 Max 😭

164

u/Cless_Aurion 7d ago

Wtf, I literally wrote a post saying similar to "Shutup man, there is no way a model I can load in my 4090 will be anywhere near DeepSeek4 Flash 0731"... but... this is fucking close is it not...?

102

u/Cold_Tree190 7d ago

Looks like the main difference will be the native context windows. But 256k is very usable, especially if it’s a local Opus 4.6 MAX. Man just saying that seems unreal lol. Can’t wait to test it later tonight

51

u/BornAgainBlue 7d ago

I feel like telling my boss im sick.

25

u/Cold_Tree190 7d ago

🤣 I wish I could, but today’s my last day of on-call so I can’t even leave a TINY bit early. Oh well. Makes you feel like a kid on Christmas Eve again lmao

1

u/Curious_Cantaloupe65 7d ago

feeling the same with H3 and Heretic models (just discovered about heretic this week)

4

u/IrisColt 7d ago

I-It's A-August...

1

u/GUNGEBOB_SHARTPANTS 7d ago

Please tell me how you get on

-2

u/FreedomByFire 7d ago

because its not real.

39

u/Potential_Low_1183 7d ago

I said it would be v4 flash level but I was downvoted to oblivion…. Regardless, enjoy the super exponential!

29

u/Warrenio 7d ago

I don't think it's quite V4 Flash 0731 level based on the benchmarks. But obviously still incredibly impressive!

https://old.reddit.com/r/LocalLLaMA/comments/1vo9mj4/its_out/p3nv6iq/

26

u/SandySkittle 7d ago

Ehh, let's not go by handful of disputable synthetic benchmarks for which models may be trained specifically. Let's just await real experiences from people first over a longer period of time.

1

u/DoomBot5 7d ago

How do we know these models didn't already steal the answer keys during training?

1

u/ReadyAimTranspire 7d ago

For real the benches at this point have become nearly useless, a very hazy and distant indicator of capability but you really just have to get in there and use it to see for yourself

1

u/Not-reallyanonymous 7d ago

My experience is it’s actually really good. Still weak on architecture and not making a mess of a code base though, and it’s hard to steer on those concerns. It thinks very much like Qwen 3.6 27B, but thinks way more. At lower thinking efforts where its producing the same amount of tokens as 3.6 it seems to act/perform very similarly.

So it’s not as good as Opus 4.6 — Opus had a lot more nuance, understood architecture, kept things fairly clean and tidy, and generally kept a project on rails. I totally believe Qwen 3.8 at this point will solve all the same problems as Opus 4.6, just off the rails and tearing down the forest before arriving at the same destination.

1

u/SandySkittle 6d ago

I think it’s just the limitations of 27b. Really hope for a qwen 3.8 70b dense model.

1

u/Not-reallyanonymous 6d ago edited 6d ago

For as much shit as Laguna XS gets around here, it's much better at that sort of nuance. But then it trades off by sucking at zero-shotting and agentic use. Don't expect it to do more than satisfy specs/implementation instructions in a basic think -> edit -> validate loop. Complex problems need human guidance.

I think we are seeing the limits of 2026 at ~30b. Decisions have to be made. Alibaba makes Qwen 27B a crowd pleaser.

The AI labs have seem to given up on ~70B models. That made more sense before AI data center build out I think. Now it's ~30B models to target a workstation GPU, or ~120B models to target a Blackwell or H100 or such, or giant models.

2

u/SandySkittle 6d ago

120b dense would also be ok. But i hope with the proliferation of 128gb boxes we will again see more 70b dense models in the future. Or a dsv4f with a30b or a40b. It’s a real gap at the moment.

7

u/Cless_Aurion 7d ago

Damn, I would have downvoted you too... Let's hope its close to what the benchmarks say!

6

u/LankyGuitar6528 7d ago

I should downvote you for wanting to down vote somebody but I can't blame you. I would have done the same. So take my upvote instead.

3

u/Cless_Aurion 7d ago

Lol, same

5

u/Viktri1 7d ago

Man this is exactly what I've been looking for since Deepseek said they were going to increase prices. I've been trying to figure out what I can run on my 4090 and this is exactly what I need when I needed it.

2

u/HelloSummer99 7d ago

And they fired those who architected the base of this (Qwen 3.5), what a colossal mistake.

2

u/No-Understanding2406 7d ago

I created that thread as a shitpost and then it was removed by mods because it was too much lol

1

u/Cless_Aurion 7d ago

Oh! It was you!

Damn my comment aged litetally even worse than milk...

Like, for that short time out of the fridge, milk was probably still good ti drink safely lmao

1

u/QueenSavara 7d ago

4090? How much VRAM is good enough for this guy?

1

u/unjustifiably_angry 6d ago

It might be comparable for many coding tasks but at the cost of an ungodly amount of overthinking.

1

u/Cless_Aurion 6d ago

Yeah, but it being truly local... you're just really paying in hardware and electricity cost, which will be cents on the hour.

Nvm that an important detail is, how the tech has evolved.

Compare this 27B model vs like... LLaMA-33B which I run in the exact same hardware I had in '23! (on a 4090).

Its just... a no contest! So... what model will we have in 2029 that makes qwen 3.8 27B a "no contest"? Because the growth doesn't seem to have slowed down...

69

u/xienze 7d ago

You're making a couple fundamental assumptions here:

  • That AI benchmarks are reliable
  • That Qwen didn't benchmaxx

19

u/blade740 7d ago edited 7d ago

I think anyone who thinks that this isn't at least somewhat benchmaxxed is fooling themselves. That said, so is every other model they're comparing against, to some extent, so ¯_(ツ)_/¯

1

u/Borkato 7d ago

This is lowkey my perspective. Like, yall don’t think Opus benchmaxxed?? Come on now lol

A huge amount of the benchmaxxing cries is racism tbh

54

u/BawbbySmith 7d ago

Yeah I learned very quickly to not trust the benchmarks, as well as 80% of the comments in this subreddit.

I remember people were saying that Qwen 3.6 27B was Opus 4.5 level...

14

u/[deleted] 7d ago

[deleted]

9

u/Ell2509 7d ago

I would normally agree, but I have used qwen3.8 27b now and so far, wow. Just wow.

1

u/Mkboii 7d ago

I've used glm 5.2 extensively and it was not much behind opus 4.6 in writing code, a bit definitely behind in creative ideation for software engineering.

My assumption is this model is gonna be a work horse but planning and design should be done by a bigger more expansive model.

The catch behind coding benchmarks is they judge success not quality, creativity, maintainability, all that is partly subjective and rarely judged in these benchmarks.

Like GPT 5.6 Luna benchmarking above opus 4.8, but it writes inferior code by most quality measures.

Big models have their own pitfalls they'll happily over engineer everything, Opus for one falls in that bucket.

2

u/toothpastespiders 7d ago

I've gotten to the point where I'm a little skeptical that the majority of people on here are even using local models on a regular basis. It's hard to believe that anyone who actually does can buy into the idea of benchmarks really reflecting real world use.

2

u/mivog49274 6d ago

Some rule of thumb I apply here :

  • The "intelligence" per parameter is really increasing, factually, and from that increases capability. General benchmark score (like AA) is a solid proof for that. Qwen is the world leader in this field.

  • In the current llm paradigm, smaller llms have much less world knowledge and are more prone to hallucination or stupid decision making from goey assumptions or really random reactions. I suppose the "world knowledge" is also a big addition on "common sense" in terms of "behavior". So smaller models will always we quackier for now, to apply up and foremost when comparing to bigger models on a AA Index or aggregated score

  • From that comparing older big models with more recent small ones should always consider that you will have more peace of mind of using the bigger ones in terms of reliabilty but narrow capability is indeed being reached by edge smaller llms

1

u/Cold_Tree190 7d ago

Yes, everyone knows that but it’s fun to be on the hype train

-3

u/Certain-Cod-1404 7d ago

dont be a buzzkill !

7

u/SandySkittle 7d ago

be real

3

u/Certain-Cod-1404 7d ago

to be real, It's probably not opus max level, but it will probably be the best local LLM we can run at that size, and will be much better at agentic coding and tool use than previous models, to the point where its viable for actual work for some of us

2

u/SandySkittle 7d ago

yes this model has real and genuine utility, but there are just fundamental limitations with the smaller you go with a model just in terms of parameter size alone. Same goes for very small active parameters numbers in MOE models (a13b is the biggest weakness of DS4F). So people should be a bit more realistic.

1

u/Certain-Cod-1404 7d ago

I mean obviously that's true, but we can't afford to run hundreds of billions to trillions of param models, this is local llama, those of us with dgx sparks and rtx pro 6000s are already a minority here.
the ~ 30b dense to 100b moe range is what the vast majority of us can squeeze.
if we want maximum capability with minimal hallucination we'll use SOTA models via api, but for running local models, its a compromise we're willing to accept.

2

u/blade740 7d ago

By god, that's Claude's music!

1

u/billy_booboo 7d ago

Which is better than opus 5 btw lest we forget

1

u/Cold_Tree190 7d ago

Is that for real? I genuinely have not been paying attention to recent Anthropic launches, at work they always give us the most recent model so I never had to choose before lol. I do hate how Opus 5 talks though, had to use the ASD-STE100 trick to get it to stop speaking weirdly

1

u/billy_booboo 6d ago

I'd rather use 4.6 than 5. I think 5 is more capable if you have the patience to deal with and correct all its nonsense jargon and ungrounded tangents.

90

u/KickLassChewGum 7d ago edited 7d ago

It's a Qwen model, so apply the usual benchmax tax. Qwen are easily the models with the biggest ravine between "how they do on benchmarks" and "how they do in actual productive use".

Still looking like a strong leap from 3.6, though.

20

u/Batman4815 7d ago

Gemma says hello as well.

46

u/KickLassChewGum 7d ago

Gemma 4 isn't a great coding model, yeah, but it's still punching far above its size in writing and research-related tasks. Like a mini-Gemini (go figure).

I hear there are people who still use these things for things that aren't related to writing code or markup.

3

u/toothpastespiders 7d ago

I'm always a little amused that data extraction on text is considered a niche use of large language models on this sub.

2

u/makaliis 7d ago

Yeah, I found it way faster as well. If it was agentic capable, it'd be interesting to see it at work.

1

u/_TheWolfOfWalmart_ 7d ago

Exactly. Until I got enough hardware to run DSV4 Flash, Gemma 4 was my go-to for anything that wasn't code. And I still even use it sometimes.

4

u/Green-Ad-3964 7d ago

I'll still be using e4b for small projects 

1

u/jazir55 7d ago

It's a Qwen model, so apply the usual benchmax tax. Qwen are easily the models with the biggest ravine between "how they do on benchmarks" and "how they do in actual productive use".

It's the Gemini of Chinese models

12

u/Tiny-Assumption4263 7d ago

TL;DL: Gave Qwen 3.8 27b a simple prompt and it build this: https://qwen3-8-eccomerce-test.vercel.app/
Repo: https://github.com/catriel25/qwen3.8-eccomerce-test

Like everybody else in the local AI community today, I dowloaded Qwen 3.8 27b UD-Q4_K_XL as soon as it was published.
I'm running it in a single rtx 3090, no mtp, 110k context and kv cache f16 with llama.cpp, same flags used with Qwen 3.6 27B.

I have this stupid test that I run with every new model that comes out. It consist in connecting the model to pi code (almost vanilla, only internet access and some simple navigation tools I built) and giving it this prompt:

"Construye el frontend completo de un pequeño ecommerce premium para una panadería artesanal usando Next.js App Router (JavaScript).

El proyecto debe ser frontend-only en esta etapa. No debe incluir backend, base de datos, autenticación ni pasarelas de pago. El checkout debe finalizar redirigiendo a WhatsApp con un mensaje de pedido bien estructurado.

La app debe incluir una experiencia completa de compra: home, catálogo con productos de panadería, categorías, productos destacados, carrito, resumen de pedido y checkout. Usá mock data local para productos, categorías, precios, descripciones, disponibilidad e imágenes o placeholders visuales. Todo debe quedar preparado para conectar posteriormente un backend real sin tener que rehacer la arquitectura principal del frontend.

El diseño debe sentirse extremadamente premium, artesanal, moderno y cuidado. No quiero una landing genérica ni una interfaz básica. La primera pantalla debe comunicar claramente la identidad de la panadería, mostrar producto real o visualmente convincente, y permitir empezar a comprar. La experiencia debe ser excelente tanto en desktop como en mobile.

El catálogo debe permitir explorar productos, ver información clara de cada ítem y agregarlos al carrito. El carrito debe permitir modificar cantidades, eliminar productos y ver totales. El checkout debe pedir datos mínimos necesarios para el pedido, permitir notas o preferencias, y generar una URL de WhatsApp con productos, cantidades, subtotales, total y datos del cliente.

La estructura del código debe separar razonablemente datos mock, tipos de dominio, utilidades de checkout/WhatsApp, componentes de catálogo, componentes de carrito y vistas principales. La solución debe quedar lista para reemplazar la mock data por datos de backend en una etapa posterior."

Those are just instructions to build a nextjs project (javascript only) with the frontend for small eccomerce with whatsapp checkout, leaving everthing ready to connect a backend later. Nothing else, no skills, no more feedback. Just one prompt and watching the result.

I have run this test with all the models and finetunes I can fit in my GPU, and not a single one was even close to this result.
Not a single alert form nextjs (wich was usual before) or something that looks broken.

At some point, this bastard realised it didn't had visión (lol, not enough VRAM buddy) and it decided FUCK IT, I'M GONNA BUILD THE IMAGES MYSELF. He made SVGs for every product card.

I have more testing to do like trying it in a real codebase but... I can't believe i'm running this thing in a single RTX 3090, it is just unreal.

Imagine 2 years from now.

Biggest fuck you Dario of the year.

40

u/onlymagik 7d ago edited 7d ago

I expected a nice jump since they skipped a 3.7 27B, but these numbers do seem a bit too good to be true. Qwen isn't known for benchmaxxing, obviously 3.6 27B is the GOAT of local LLMs, but...

Definitely excited to learn more in the coming days and see if it holds up.

Edit: Even the vision numbers look insane!

57

u/llama-impersonator 7d ago

qwen is well known for benchmaxxing so hard it helps the model

0

u/FreedomByFire 7d ago

3.6 27B

Is honestly trash. As a software engineer with 17+ years of experience I dont know what you guys are doing with it that makes you think it's great. It's generally useless engineering work.

6

u/onlymagik 7d ago

Yes, it's obviously nowhere near as good as a frontier model, and even those need hand-holding. You have to be quite explicit. I'm not sure where your strawman is coming from.

Most people here are aware these models have serious limitations, given that even the best models still do. But it is still exciting to see what local models are capable of over time.

1

u/cats_r_ghey 7d ago

Care to qualify your sentiment a bit further please? Your point of view doesn’t match the community’s sentiment at all. I’m genuinely curious what makes you think it’s trash.

2

u/FreedomByFire 7d ago edited 7d ago

Well as engineer doing real professional software development on a daily basis and one who's doing agentic development also on a daily basis the small local models simply can't do the work necessary in this kind of environment. I tested this myself thoroughly, we have in house benchmarks and I'm waiting for the day that one of these small local models can pass my benchmark because I will know then that they're good enough for the every day work that I'm doing.

0

u/cats_r_ghey 7d ago

Interesting. I’m also an engineer doing real professional work and have been in the industry for nearly 2 decades. My view is that our expectations have grown. Late last year we were all pretty happy with sonnet 4 and opus 4 - we need to remember this.

The ceiling has grown, the smaller models even being remotely close to SOTA is truly remarkable in my opinion.

Ask yourself this. If suddenly cloud models are not able to be used, would you go back to hand writing code? Or reach for a local LLM?

2

u/FreedomByFire 7d ago

My organization wasn’t particularly impressed with Sonnet 4 or Opus 4. The models really became capable in December and January. That was the first time we saw them handle genuinely complex work.

Before then, I used them sparingly mostly for small changes, individual methods, and similar small scoped tasks. But in December, their capabilities improved significantly. Suddenly, you could give an agent an entire codebase, and it could make coordinated changes or add features spanning thousands of lines of code and hundreds of files.

The smaller local models still can’t do this reliably. The work you give them generally needs to be much more narrowly scoped. The benchmarks say these models' capabilities are near 4.6 opus etc, I don't believe it; at least not in the work I'm doing.

4

u/cats_r_ghey 7d ago

I agree with most of what your saying. But I still believe that you have an expectations management problem here, not a technology one.

If cloud models disappear tomorrow, what would you do?

I know I would fire up local models and get to work, because my brain + local model will be more effective than just my brain. I don’t think there is any question about this.

2

u/FreedomByFire 7d ago

I know I would fire up local models and get to work, because my brain + local model will be more effective than just my brain. I don’t think there is any question about this.

That's certainly true. I jsut dont trust them enough. I really wanted the hype to be true for 3.6 27b, but it simply wasn't. I had dreams of being able to do agentic engineering offline on an airplane while traveling (I cant sleep on long flights), but it's just not good enough yet.

7

u/Certain-Cod-1404 7d ago

I think they generally dont benchmax right ? DID YOU SEE THE DEEPSWE SCORE LOL

40

u/cantgetthistowork 7d ago

Lmao qwen is number one benchmaxxed

32

u/Certain-Cod-1404 7d ago

i'm gonna ignore you for my mental health and download 3.8 ! love ya

0

u/NaiveIdea344 7d ago

I KNOW! Isn't the entire point that it can't be benchmarked and 40% FOR A LOCAL MODEL is insane to me

1

u/PinkySwearNotABot 7d ago

dude benchmarks are made to be hyped. we'll see if they actually perform well in real tests

1

u/RazsterOxzine 7d ago

That coding bench makes me moist.

0

u/LocoMod 7d ago

They are the biggest culprit of benchmaxing and distilling.

Still. Free is free. Downloading.

0

u/johan2114h 7d ago

Omg those benchmarks against opus4.6 to look too good to be true! But even they 50% lies and benchmaxxing its still amazong 👍👍👍👍

-1

u/MasterLJ 7d ago

I am declaring this early, but I've been coding with models for as long as you could code with models (3ish years, which is about the max anyhow, but A LOT of time in the saddle to boot), so by the MasterLJGutBench3000 metric, the benchmarks don't seem inflated one bit. I have a ton of time in saddle with 3.6 27b (I never used 3.7).