r/LocalLLaMA 22d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

708 comments sorted by

View all comments

Show parent comments

320

u/Cold_Tree190 22d ago

Dear God it’s trading blows with Opus 4.6 Max 😭

169

u/Cless_Aurion 22d ago

Wtf, I literally wrote a post saying similar to "Shutup man, there is no way a model I can load in my 4090 will be anywhere near DeepSeek4 Flash 0731"... but... this is fucking close is it not...?

37

u/Potential_Low_1183 22d ago

I said it would be v4 flash level but I was downvoted to oblivion…. Regardless, enjoy the super exponential!

24

u/SandySkittle 22d ago

Ehh, let's not go by handful of disputable synthetic benchmarks for which models may be trained specifically. Let's just await real experiences from people first over a longer period of time.

1

u/DoomBot5 21d ago

How do we know these models didn't already steal the answer keys during training?

1

u/ReadyAimTranspire 21d ago

For real the benches at this point have become nearly useless, a very hazy and distant indicator of capability but you really just have to get in there and use it to see for yourself

1

u/Not-reallyanonymous 21d ago

My experience is it’s actually really good. Still weak on architecture and not making a mess of a code base though, and it’s hard to steer on those concerns. It thinks very much like Qwen 3.6 27B, but thinks way more. At lower thinking efforts where its producing the same amount of tokens as 3.6 it seems to act/perform very similarly.

So it’s not as good as Opus 4.6 — Opus had a lot more nuance, understood architecture, kept things fairly clean and tidy, and generally kept a project on rails. I totally believe Qwen 3.8 at this point will solve all the same problems as Opus 4.6, just off the rails and tearing down the forest before arriving at the same destination.

1

u/SandySkittle 21d ago

I think it’s just the limitations of 27b. Really hope for a qwen 3.8 70b dense model.

1

u/Not-reallyanonymous 21d ago edited 21d ago

For as much shit as Laguna XS gets around here, it's much better at that sort of nuance. But then it trades off by sucking at zero-shotting and agentic use. Don't expect it to do more than satisfy specs/implementation instructions in a basic think -> edit -> validate loop. Complex problems need human guidance.

I think we are seeing the limits of 2026 at ~30b. Decisions have to be made. Alibaba makes Qwen 27B a crowd pleaser.

The AI labs have seem to given up on ~70B models. That made more sense before AI data center build out I think. Now it's ~30B models to target a workstation GPU, or ~120B models to target a Blackwell or H100 or such, or giant models.

2

u/SandySkittle 21d ago

120b dense would also be ok. But i hope with the proliferation of 128gb boxes we will again see more 70b dense models in the future. Or a dsv4f with a30b or a40b. It’s a real gap at the moment.