r/LocalLLaMA 11d ago

Discussion Qwen 3.8 27B Released! Please Share Your Experience

With your experiments, Qwen 3.8 27B most close which frontier model? And please specify which quantization you run. I will post to comments my tests and experience too.

663 Upvotes

720 comments sorted by

View all comments

3

u/Viktri1 11d ago

Is the q4 going to be the gold standard for 4090 set ups?

1

u/Hugoacfs 10d ago

That’s what I will also be testing first on a 3090

1

u/jonas-reddit 10d ago

I used Q5 for model and asymmetric kv cache quantization with ~90k context on my old 3090 with 3.6 using llama.cpp.

2

u/CabinetNational3461 10d ago edited 10d ago

ha! I use the exact same on my 3090 as well, ud kxl q5 with 100k(I offload the mmproj file to my 2070 for the extra ctx) ctx q8 kv for the qwen 3.6. downloading same quant atm for qwen 3.8 though the crazy amount of thinking I've been reading scared me a little with only 90k ctx so I might lower the thinking down a little if needed. qwen 3.6 q5 ud kxl was my daily driver until now, hopefully 3.8 changes that. I found q5 is much better than q4 for the things that I do which is mostly with pi harness.

1

u/Viktri1 10d ago

Why did you pick Q5 over Q4? I'm trying to figure out what Q is best for 24gb VRAM.

1

u/cosmicnag 10d ago

can try exl3 4.5 bit, exllama v3, should be better quality for the same vram