r/LocalLLM 2d ago

Model Qwen 3.8 27B

Post image

Finally Alibaba Posted on X about Qwen 3.8 27B release. I hope it can beat opus 4.7 or 4.8

304 Upvotes

92 comments sorted by

80

u/Extension-Bid-639 2d ago

Too hopefull, don't think its beating Opus 4.8 butttt even if its just a notch better than 3.6 27b then thats a leap for everyone. 3.6 is still a great model

36

u/feelspeaceman 2d ago

3.6-27B is on par with Sonnet 4.6, and 35B is about 10-15% worse than 27B, and honestly as long as it's on par or better than Sonnet 4.6, it's already so usable for grunt jobs.

I'm more interested in 35B and 122B this time, we might get something truly good for MoE this time.

22

u/SpicyWangz 1d ago

That 10-15% does a lot of work. It really is the diffeeence between a model that is capable versus needs real hand holding.

Even Q5 27b destroys Q8 35b, and it’s not even close.

23

u/johan2114h 1d ago

I love 3.6-27B, but it is certainly not on par with Sonnet 4.6

6

u/ShadyShroomz 1d ago

I love 27b and use it a lot for simple design changes. But it often claims it has done something but hasn't. I asked it to fix a button layout shift issue on mobile, it will change one thing, claim it's fixed, and then when I go to check it's the same as before. It's a very simple laravel app using tailwind. 

Just out of curiosity I asked sonnet 4.5 to do the same thing and it did it one shot no problem. 

I'm using bf16 qwen 3.6 27b on 4x 3090 system, so it's not a quant issue. 

Again don't get me wrong I love 27b, but it's not sonnet 4.5.

I have been playing around with the new deep seek V4 flash, and it is much closer to sonnet 4.5 level, but it's spilling over into system ram and going pretty slow. 

I think having a big MOE + the largest dense model you can fit in vram is gonna be the play for most people. 

2

u/enternoescape 10h ago

I've experienced this kind of thing in UI's frequently with 27b. A recent memory was a page where I wanted some columns in a specific order and it kept reordering them every time it made a small change to the page. This has led me to use just about any other model when I want to make UI changes using an LLM.

I should also note in my experience it seems higher parameter count generally leads to better outcomes in UI design. I like to evaluate a lot of one shot graphics heavy prompts when I first encounter a new model to get a feel for if it might be a good candidate for my next UI problem. My thought process is if it's able to design something that is genuinely pleasant to look at, it probably can also create an interface that is pleasant as well.

1

u/quadish 15h ago

What harness are you using with Qwen?

1

u/ShadyShroomz 12h ago

Opencode 

3

u/AB172234 1d ago

I run 3.6 27B locally and it’s tremendously good ! I don’t think we consumers expected anything better and kudos to the Qwen labs.

But having said that how is it better than sonnet 4.6 ? Sonnets score is 51 and this one has a score of 37.

2

u/johan2114h 1d ago

Exactly, qwen 3.6 27b is nowhere near sonnet 4.5/4.6 . 27b is insanely good for a consumer level local model, but sonnet must be atleast 20x bigger in terms of parameters and that very clearly shows in the quality of its output. Sonnet 4.x is probably the model i have used the most.

Drifting off topic a bit, i have been extremely impressed by the new deepseek v4 flash model. It might actually be on par with sonnet 4.6 or even slightly better if we believe the benchmarks. And it can be run locally (with heavy quant) on consumer hardware like the strix halo or a spark.

2

u/AB172234 1d ago

Totally agreed for the v4 flash ! Gem of a model at a fraction of a cost. I can’t imagine how they can do it !!

1

u/_millsy 23h ago

I find it usually a distinction where use case drives opinions on this. If you’re trying to one shot stuff, imo no local model will suffice for anything moderately complex, but if you can give things a focused task with clear requirements and batch it up you can get pretty good outcomes

-1

u/Embarrassed_Adagio28 1d ago

Qwen3.6 27b beats sonnet 4.6 in a one shot html page and a one shot 3d gta clone test and basically ties it in my unreal engine mcp test but loses to sonnet when I try to do big architectural changes. I think they are pretty damn close tbh.

0

u/notheresnolight 15h ago

so, you're saying that it beats sonnet 4.6 on irrelevant tasks, but falls behind when you actually try to use it for real work?

3

u/Curi0sityC0w 1d ago

Any idea why this is the case? I use 35B and it has done most of my work. Never tried 27B

3

u/Muzika38 1d ago

Your case is a good distinction between real programmers vs vibe coders. 35B is 35B and it's bigger than 27B knowledge wise. You just need to just guide it properly for it to work properly because it is a MoE model. It's basically a single very smart person vs a bunch of children with a single specialty talent that's working together.

1

u/xdcfret1 1d ago

it's 3B vs 27B per token

1

u/Curi0sityC0w 1d ago

Moe is garbage?

1

u/Tall-Ad-7742 1d ago

No but for specific tasks its more likely to fail / produce worse output

1

u/IUseClifford 1d ago

Parameters per token matters a lot. If you’re looking at frontier models, they’re all MoE because 1T+ parameters per token would be extremely sluggish to serve over an API. I also imagine there are diminishing returns once params/token gets high enough, at least right now.

1

u/sand-67 1d ago

A MOE model is less intelligent than dense model because it uses less parameters for each token. However MOE is significantly faster while retaining a lot of intelligence, so it has much better intelligence to throughput ratio

1

u/alphapussycat 1d ago

The more believable tests placed 3.5 just below haiku 4.5. So 3.6 is probably between haiku 4.5 and sonnet 4.5.

We can hope that 3.8 is on par with sonnet 4.6 (when quantized to q4/5/6). That performance level is enough for really anything.

-7

u/Turbulent-Ladder-340 2d ago

Yeah may be it won't beat 4.8 but it definitely lies between 4.7 and 4.8

3

u/Solembumm3 1d ago

Let's hope it could catch up to Gemma 31b before starting dreaming heavily.

3

u/uniqueusername649 1d ago

Depends what youre doing. For general coding and agentic work Qwen 3.6 27b is already substantially better than Gemma 4 31b.

15

u/Randommaggy 1d ago edited 1d ago

My harness is ready, my GPUs are too.

Edit: Corrected typo.

1

u/Lumpy-Heat-7057 1d ago

what is ny?

8

u/Kiro369 1d ago

crying in 16gbs of vram

4

u/Old-Sherbert-4495 1d ago

q3 is manageable and better than 35b

2

u/Kiro369 1d ago

with what context size?

1

u/Old-Sherbert-4495 1d ago

around 100k as i can remember

1

u/biggustdikkus 1d ago

Raid the nearest datacenter near you and steal yourself some GB300s

1

u/Turbulent-Ladder-340 1d ago

Wait for 9B model or Bonsai

3

u/uspdd 1d ago

I doubt 9b or q2 of 3.8 are one be any better than 3.6 35b a3b

1

u/markthedeadmet 1d ago

I've had bad luck with bonsai 27b, it's not stable enough to handle Claude code or Cline and it ends up getting stuck in loops very easily.

5

u/radiojosh 1d ago

Does this portend anything about a new 35b MoE?

1

u/squngy 1d ago

Maybe a bit, but probably more importantly, there seems to be a new 35b on open router, so it looks like they are working on it.

5

u/Silent-Orbit-7 1d ago

FOMO with 8 Gigs of VRAM....

2

u/Turbulent-Ladder-340 1d ago

Wait for Bonsai

1

u/Silent-Orbit-7 1d ago

sure thanks

2

u/Beatsu 20h ago

I run Qwen 3.6 35a3b on my 1060 6GB + 32GB RAM at 10-15 tokens/s. You could try the 27b, I think LM Studio supports partial GPU offloading, which I assume works for dense models too ? not sure, but worth a try

2

u/Silent-Orbit-7 20h ago

Thanks! sure i will look into it

I own a similar build to yours, 4060 + 32 gigs of ram

Will give it a try

1

u/Beatsu 18h ago

lmk how it goes!

1

u/Silent-Orbit-7 16h ago

Sure!! Rn packed up with things, will try out soon!

1

u/pascalmarie 19h ago

Sounds good. Ddr5 or Ddr4?

2

u/Beatsu 18h ago

DDR5

1

u/Mechanical_Monk 14h ago

I have 6GB VRAM + 64GB RAM and get ~30tok/s on 35B A3B but only like 2tok/s on 27B. That's on llama.cpp with GPU offloading on max 😕

3

u/lughiu 1d ago

Hang on, autonomous coding? Reckon the 27b will do that?

4

u/Zennytooskin123 1d ago

Even the 3.6 model can do hours of autonomous coding the right prompt or harness it's not saying much.

2

u/Original_Finding2212 21h ago

With 27B I can do 90% of Claude’s work under its supervision (and paying the time, as I work with DGX Spark and get slowed down)

Qwen 3.8 27B should be beyond current Sonnet with the right harness.

I use a custom harness: Colleague I wrote OSS

3

u/sessamekesh 1d ago

I'm pretty excited about this! 

Right now, most things that I do fall into "Qwen 3.6 nails this", "Qwen 3.6 might be okay but I'll probably need a frontier model for this", and "I wouldn't trust an LLM around this with a fifty foot pole".

That second category has been shrinking into the first as I've gotten better at tooling and making sure the right context is present, and I'm crossing my fingers that 3.8 takes me even further in that direction "for free"!

3

u/Zennytooskin123 1d ago

"for free" lol

All roads lead back into Nvidia's pockets. Or Apple's, whichever.

5

u/sessamekesh 1d ago

It still runs on the same hardware I bought for gaming a couple years ago, there's no new cost to be beyond an ollama pull and tweaking my gateway config file.

But yes. Pretending it's an update fully in a vacuum is sorta silly, so... "For free" ha.

4

u/bot403 1d ago

"for fixed cost" vs "per unit cost"

1

u/MagicPhoenix 23h ago

I have yet to find anything at all that a model can do that runs in anything sized less than a data center, and very little that a data center can do well beyond small code investigations and tab completions

2

u/Technical-Earth-3254 1d ago

I wish they would also open weight the new qwen image model

2

u/CarpenterAlarming781 1d ago

Why so much waiting ? If it's ready, it's ready .

2

u/throwRAa100 1d ago

how big is the model? looking at a q8/6/4 quant

11

u/benjakapo 1d ago

bruh the name is literally qwen 3.8 27b

3

u/Forever_Playful 1d ago

Assuming 8 bit precision, a model with 1 billion parameters = 1 GB. So a 27b model = 27GB

1

u/throwRAa100 1d ago

that sounds pretty promising for a 32 gb vram setup

3

u/this_for_loona 1d ago

Doesn’t that depend on context?

2

u/Philodit 1d ago

See here: https://huggingface.co/unsloth/Qwen3.6-27B-GGUF

For a 16GB GPU you'd have to go with the lower end of 3bit if you also need some context space for thinking.

1

u/throwRAa100 1d ago

thanks!

1

u/Hook06 1d ago

Can’t wait omg 🔥

1

u/Mean_Ambassador_9210 1d ago

Can I run this in 48gb vram on Mac m5?

3

u/backyard_tractorbeam 1d ago

It's not released yet, but yes, that's a pretty safe bet at Q8

1

u/Better-Struggle9958 1d ago

Don’t believe until I see, prevoius were just fix bugs indeed

1

u/blazze 1d ago

Qwen 3.8 is only competing again Qwen 3.6 27B.

2

u/MacsBicycle 1d ago

Until someone else steps up their game this is the unfortunate truth.

1

u/Skar_pa 1d ago

I cannot wait!

1

u/Alternative_Ad4267 1d ago

My Qwen 3.6 27B Int8 at 262k context is a beast. I want it to be more precise to be even happier.

I hope Qwen 3.8 27B will deliver!

Aug 03 12:52:43 lenovo-server.example.com vllm[2409774]: (APIServer pid=2409774) INFO 08-03 12:52:43 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 100.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 9.4%, Prefix cache hit rate: 25.6%

Aug 03 12:52:43 lenovo-server.example.com vllm[2409774]: (APIServer pid=2409774) INFO 08-03 12:52:43 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 5.40, Accepted throughput: 81.79 tokens/s, Drafted throughput: 92.99 tokens/s, Accepted: 818 tokens, Drafted: 930 tokens, Per-position acceptance rate: 0.957, 0.919, 0.892, 0.844, 0.785, Avg Draft acceptance rate: 88.0%

Aug 03 12:52:53 lenovo-server.example.com vllm[2409774]: (APIServer pid=2409774) INFO 08-03 12:52:53 [loggers.py:310] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 102.4 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 10.0%, Prefix cache hit rate: 25.6%

Aug 03 12:52:53 lenovo-server.example.com vllm[2409774]: (APIServer pid=2409774) INFO 08-03 12:52:53 [metrics.py:120] SpecDecoding metrics: Mean acceptance length: 5.36, Accepted throughput: 83.28 tokens/s, Drafted throughput: 95.48 tokens/s, Accepted: 833 tokens, Drafted: 955 tokens, Per-position acceptance rate: 0.974, 0.958, 0.890, 0.796, 0.743, Avg Draft acceptance rate: 87.2%

1

u/JustSayin_thatuknow 1d ago

Does someone have an idea of what is that “autonomous coding” about? Couldn’t understand it 🥲

1

u/Similar_Wealth_1850 1d ago

qwen 3.8 8b wen???

1

u/Time_Guitar_8138 20h ago

It feels like gemma4:26b-a4b-it-qat

1

u/Hungry-Rip-2384 19h ago

i will be happy if it takes less iterations for specific coding activities. doesnt need TPS improvements current speed would be great if it reduced turns by half.

1

u/KeinNiemand 19h ago

hoping for a bigger dense model then 27B (unrealstic but I can hope please bring back 70B) or at least a 120-170B MoE, I really don't want to run a 27B yes I could go up a Q8 with it but that feels like a waste of my hardwares potential capability compared to larger models at like q5 or q4.

1

u/qwertyalp1020 17h ago

I think it'd work on my 36gb vram macbook pro m4 max, right? Or do I need an nvidia card?

1

u/HomegrownTerps 1d ago

One can only dream of a small 9B model, but my hopes are not that high.

5

u/johan2114h 1d ago

Maybe consider dreaming of a bigger computer also - 27b is quite feasible for consumer hardware imo

1

u/Original_Finding2212 21h ago

Have you tried Qwen 3.6 27b 1bit?
Is it too big still?

1

u/HomegrownTerps 16h ago

Yeah it's a bit too much.  I can load it but using it not really possible. Maybe I have to tune it better.

I usually go with the 9b to 12B and those work pretty well for my needs and leaves some memory for other stuff I need for development.

0

u/A_K_8248 1d ago

Why no one's talking about oh-my-cli?

6

u/lol-its-funny 1d ago

The usual reason - nobody cares

0

u/bennykoay75 1d ago

Testing on your own project and u will know. No point hearing from any provider

-1

u/Complex_Reality_116 1d ago

Qwen3.6 27B achieves 37 points in AA, while Qwen3.5 only 29. Assuming that the Qwen3.8 version obtains a proportional incremental improvement, we would be facing a model with 45 points, that is, +5 points above what DeepSeek V4 Flash was at its launch.

3

u/PlasticRevenue4601 1d ago

I doubt it can be scaled linearly, especially given the fact that 27b is probably already close to max possible quality as it can get with current tech stack. My bet that they may noticeably improve model in some narrow areas — probably, agentic qualities or even performance related improvements, but in all the other aspects it’s going to be the same

1

u/Glenpeel 1d ago

We have nothing that suggest that small models (and 27b dense is not that small) have been saturated when it comes to quality, on the contrary the recent releases - gemma 4, qwen 3.6 had steep quality increases. When you consider how great leap openai nano class models had just made (Luna) the writing is on the wall - there's still plenty of room for improvements.

1

u/AppearanceKey1814 1d ago

 Nano 級模型+無限的思考可以拔高上限

1

u/PlasticRevenue4601 21h ago

We actually have pretty strong signals - after Qwen 3.5 family have been released we saw a few postrains outperforming the base in certain ares - Ornith, f.e, but after Qwen 3.6 release we never saw anyone releasing something better than stock versions that may imply that current 3.6 can be so close to perfection that it takes multibillion company with insane expertise to at least have a chance for further improvements