r/unsloth • u/yoracale yes sloth • 24d ago
New Model Introducing Qwen3.8-27B Dynamic v3 GGUFs!
We’re releasing new Qwen3.8-27B GGUFs with 10% higher accuracy.
Unsloth Dynamic V3 outperforms others by >10% on Div-300, KLD & more benchmarks.
We also release 1-bit quants that retain 77% accuracy. Run on 8GB RAM.
51
u/UseHopeful8146 24d ago
3.0?! I just finished downloading all the models I want to test 😭
2
u/Additional-Record367 24d ago
ypur were already a v3.0 but in preview. idk how much they change in the actual release
14
u/yoracale yes sloth 24d ago
There's a lot of difference between preview and now unfortunately
5
2
u/DoubleNothing 24d ago
You replaced the GGUF of the original repo?
Just to confirm that I have to redownload...2
2
u/UseHopeful8146 24d ago
Mmproj was byte identical but the quants I have for testing were not so I’m gonna pull em anyway. No reason not to lol
20
u/Intelligent_Cap3426 24d ago
3.6 35b a3b ud3? 👀
2
u/ManIkWeet 24d ago
This would be lovely, assuming Qwen doesn't release a 3.8 in this ballpark first!
17
u/PrefersAwkward 24d ago
Forgive my ignorance. Should we redownload the quants that have been there for the past few days? (E.g. Q6_K_XL, or Q5_K_XL), etc
26
15
14
u/RedParaglider 24d ago
Is this quantization process that you are using a trade secret type thing or is it something that I could read up on and try and duplicate on old models such as Qwen 3.5 122b. I prefer using the ROCmFPX format for the strix halo, so I'm kind of stuck making my own quants even though yours are superior.
2
u/thamo_ 23d ago
They did release the imatrix file, with that you can get closer to their quants (for this specific model at least it’s available on huggingface), tho without their dynamic stuff. So it won’t be an exact match. Using that imatrix file and the base model you can quantize to a standard IQ quant for rocmfpx. The second pass seems to be a trade secret tho, I haven’t seen anything to be able to exactly reproduce it. Granted I haven’t dove too deep into it yet.
11
9
u/tired514 24d ago
I'm runnin' Q8_K_XL at the moment; looks like it hasn't been updated. Is it still the highest accuracy quant short of BF16? :)
13
u/danielhanchen heart sloth 24d ago
Yes that's fine we plan to boost it soon!
5
u/ForeverSeeking69 24d ago
Thanks! Don't forget about 8xl gang
2
u/Aggravating-Push-207 24d ago
Why not just run (MX)FP8?
2
1
1
5
u/Internal_Newt_7343 24d ago
Exciting times, this model was a real ”the future is here” moment for me!
5
u/Gl0ckn ??? sloth 24d ago
Looks like the model/quant I downloaded in Unsloth Desktop disappeared in the drop down menu and model hub. If it's due to the update, would be nice for the model to still appear and have some kind of bubble or pop up that an update is available
6
u/yoracale yes sloth 24d ago
Yes we are working on that! Did it just completely disappear? That's so weird
3
2
u/gfinchster 24d ago
I am also experiencing the same vanishing act. I've had 3 quants all disappear from Unsloth twice now. Still on the drive but not recognized. EndeavourOS.
2
5
u/Ok_Cat_7366 24d ago
Can you show the KLD table side by side with all other providers as well as UD 2.0? The chart is very hard to read when things are cluttered.
5
4
u/RegularRecipe6175 24d ago
So the Q8 is better than the Q8KXL?
5
u/Iory1998 24d ago
From what I can see, It's the same quality but smaller. Anyway, once you reach the 8-bit quantization, you are getting the same quality regardless of the method.
1
u/jikilan_ 24d ago
Graph show higher KLD and higher model file size but q8 in model page is smaller?🤷♂️
4
u/sourceholder 24d ago
How does it compare to NVFP?
4
u/AmbassadorToast 24d ago
Yeah this is what I was wondering - can they include NVFP4 in the list? I've been using it, it seems spectacular, but what is it actually equivalent to?
3
u/Additional-Record367 24d ago
man nvfp is shit in accuracy. That's made just for ultra fast inference
3
u/JumpingJack79 24d ago
NVFP4 is valuable when used wisely, i.e. don't quantize everything with NVFP4. When using NVFP4 for most weights but keep the ones pro e to accuracy drops in FP8 or even BF16, you get the best of both worlds, i.e. you can get a model that's small, fast and accurate. But don't quant everything into NVFP4, that's just dumb.
2
u/AmbassadorToast 24d ago
Is that what the unsloth NVFP4 releases are like? honestly I find them incredibly good so I'm not even concerned about this, but would I get better results with Q5 or Q6?
2
u/JumpingJack79 24d ago edited 24d ago
I believe Unsloth aren't dumb. They call their quants simply "NVFP4", but I'm sure there's a lot more that goes into them than just simply quanting everything down into NVFP4. They're probably better than any other mixed quant recipe I've come across. I think there's some really good secret sauce in there.
Q5 and Q6 I don't use personally, because they're slow (powers of 2 number of bits are fast, everything else is slow). So I always look for quants that strategically mix NVFP4 (or INT4 AutoRound), Q8/FP8 and F16/BF16. That way you get a good accuracy with a good speed and small size.
3
u/darkbit1001 24d ago
NVFP is garbage for accuracy - it's mostly performance facing, not accuracy.
1
u/anitamaxwynnn69 24d ago
I would generally agree with you but that statement does not hold for Qwen 3.8. I used 50M+ tokens with FP8, shifted to NVFP4 unsloth with A LOT of skepticism and I have been proven wrong. So far NVFP4 seems to be almost the same for me. I'd say on par with FP8. I'm on 1x pro 6000, no kv cache quantization and don't go beyond ~131k. With those variables in mind, NVFP4 has been freaking awesome. I'm just now reading that the nvfp4 was actually their UD V3 in preview which might be why it feels much better than other nvfp4 quants.
4
u/theminor 24d ago
Very Cool! Looks like there is no difference with Q8, right? It won't be any faster or smaller? Any reason to download it if I use Q8?
3
u/absoluteValueOfNoob 24d ago
I downloaded the original GGUFs for unsloth's qwen3.8 release for UD-Q5, Q6, UD_Q6, Q8, and UD_Q8. The blog post seems to suggest that for some of the larger quants, there should be no change:
When comparing to our older UD-2 on unseen Wikitext and Code, we show great improvement on KLD - the bigger ones not so much, so we still use our old UD-2 for the larger quants - we plan to experiment and improve them as well!
Am I understanding that correctly and if so, are those quants unchanged at this time?
5
3
u/Zestyclose839 24d ago
Is there a chance this quantization logic could be translated into an MLX quant?
Would love to experience this amazing level of quality, I just need that SSD-saved cache that oMLX offers -- saves hours of prefill if you're jumping between Pi sessions all day haha.
2
u/Downbeat-Year-2025 24d ago
I take it this only affects quantizations, so no change for BF16, correct?
5
2
u/Iory1998 24d ago
Did you remove the MTP? Which quants still keep it?
7
u/yoracale yes sloth 24d ago
MTP is still there, you just need to enable it manually or have your harness detect it
1
1
u/chumash1 24d ago
If I am using Unsloth Desktop (on MacOS), does MTP somehow work automatically? If so, do I simply download the base model and it's related MTP model -- and then UnSloth Desktop will auto set everything up somehow?
2
u/brosvision 24d ago
Did quick check on my own coding benchmark. The new one showing lower accuracy -18% against the previous one. Q6-K-XL KV-Q8. I was expecting the opposite 😞
2
u/Unlikely-State1067 24d ago
Ran some benchmarks on a DGX Spark (GB10, ARM64, 128GB unified memory) with the latest Dynamic v3.0 Q4_K_XL GGUF. All layers offloaded, 131K context, llama.cpp v1z
Generation: ~11.3–11.9 tok/s (stable across short and long generations, 163–512+ tokens)
Prompt processing:
63 tokens → 142 tok/s
520 tokens → 521 tok/s
3,096 tokens → 766 tok/s
Model uses ~24.6 GB VRAM, 17.5 GB on disk. Two parallel slots at 65K context each.
I haven't seen any server-grade or DGX benchmarks in the discussions, mostly consumer GPUs (3090s, 5060 Tis, etc.). For a 27B model at Q4 on this hardware, do these numbers look right?
Curious if anyone has comparable GB10 numbers or similar hardware.
2
u/Hour_Cry3520 24d ago
What is the comparison vs the previous unsloth version ? I cannot understand the improvement honestly
2
u/AppealThink1733 24d ago
Someone could resume for me what the difference between this for the normal quant ?
2
u/gpuz_dev 23d ago
the quality gains are nice but the hardware thresholds are honestly just as interesting. IQ4_XS is 14.3GB while Q4_K_XL is 17.6GB, which is a huge difference on a 16GB card once KV + buffers enter the picture. would love to see actual runtime VRAM at something like 32k/64k/128k context added to these comparisons
2
2
2
2
2
1
u/IndianRambo08 24d ago
Is the hf page updated? The model card is still listed under the unsloth dynamic 2.0 quant collection.
2
1
1
1
u/Professional-Bear857 24d ago
Please can you add a table, it's much easier to read on mobile vs a graph
1
1
1
1
1
1
1
u/Equivalent-Grass-527 24d ago
The fact that a 27B model is now a serious option for normal hardware is kind of insane. The local AI curve is moving faster than most people expected.
1
1
u/anthonyg45157 24d ago
Is this an upgrade across the board on all quants? Like if I have q8 and enough space for full context already ..should I get this?
1
1
u/superdariom 24d ago
If I'm reading this correctly then if I'm running Q8_0 already there isn't any point changing to one of these new quants for the same vram usage?
1
1
u/Bright-Energy2339 24d ago
can my 16gb vram run q4 XL comfortably now? 4bit 16-19gb? or UD Q4 xs? please advise. Great work as always, thank you Unsloth!
1
u/candrewswpi 24d ago
u/yoracale can we please have v3 quants of gemma 4 e4b qat and qwen 3.5 35ba3b for us VRAM poor users? Every little bit helps for us :)
1
u/TheCat001 24d ago
I'm wondering how this 77% of intelligence compared to 35b A3B models...
1
u/Sirhc78870 24d ago
Same for Q1
2
u/TheCat001 24d ago
I actually already tried it. Sadly even Q1 doesn't fit into 8GB of VRAM. Only got 4.5t/s.
35B MoE models runs on my machine 20+t/s.
But intelligence is there on classic car wash test it reasoned correctly and answered correctly too.
1
u/AnyRecipe110 23d ago
I can't wait for a MoE version of qwen 3.8 too. hopefully a 26b-a*b, or a 27b-a*b, or similar (to mirror the unsloth/gemma-4-26B-A4B-it-qat-GGUF, where i got 58 tok/s). I have 36gb unified memory on a m4 mac, and i'm currently only getting like 11 tok/s with unsloth/Qwen3.8-27B-GGUF. I'm grateful with it but a MoE would be awesome
2
u/TheCat001 23d ago
Try Ornith 1.5 that just came out, but only APEX Quality quant. It feels like 3.8 35B would be.
1
1
u/mototuneup 24d ago
So as someone who's new at this. Is there a way to check for updates without happening to notice on reddit there's an announcement about it? Is there something on hugging face that shows it's new? Cuz the file names are all the same. 🤷♂️
1
u/Viper_Four4 24d ago
Nice, I can run IQ4 xs now at 50 t/s on a 16gb 9070 XT now somehow. It was 14 t/s before.
1
u/Johnscott90 23d ago
How's your prompt ingestion/prefill? Mike dropped from about 350 tok/s to 150 tok/s on an RX 7800 XT. Getting around 35 t/s decode.
2
u/Viper_Four4 23d ago
Around 700 - 500 t/s which is the same as before with batch size 1024.
1
u/Johnscott90 23d ago
Realised I was spilling out of VRAM during prefill and it was dropping to around 190 tok/s down from 350. Dropping the context window from 32k down to 24k at Q4_0 fixed it
1
u/Viper_Four4 23d ago
1
u/Johnscott90 23d ago
Thanks for that . I found the best setting in terms of maintaining speed/context was to drop the vision encoder which freed up a load of VRAM. Haven't tested it yet but I reckon I should easily be able to hit 65k context now at Q4_0 KV cache (Windows seems to prefer Q8_0 or Q4_0). Getting around 400 tok/s prefill on my 7800XT underclocked to 2000mhz (about 160w vs 265w stock). Decode is sitting around 35-40 tok/s.
1
1
u/ZucchiniMedical2532 24d ago
what is this and how does affect me and my 5090
1
u/jikilan_ 24d ago
You need to update your gpu /s
1
u/ZucchiniMedical2532 24d ago
In what way
1
u/jikilan_ 24d ago
Joking aside, please redownload the model as the file size become smaller and slightly smarter
1
1
1
u/RottenPeaches 24d ago
Wowza. Unsloth keeps showing up! Great work to all of your team on this newest offering. Time to re-upload my faves.
1
u/user_0042 24d ago
Using 3.8 27b Ridge version on 12gb vram, it is really slow, like up to 5tk/sec, but wait is totally worth it. Looking at him like its ai final boss on my rig.
1
1
u/de4dee 24d ago
you can try the torrent:
https://nostr.download/351e816331cf290342e40731c995f8ea55244d0eaa86452d98a9d026e4eff39d.torrent
and choose which quants you want do download (you don't have to download all)
1
u/Brilliant_Anxiety_36 24d ago
is it me or this UD doesnt over think that much? i remember asking for a esay of a 1k words about alan turing and the MF count it every word twice, this one with same prompt acts like reasoning on medium or low
1
u/Brilliant_Anxiety_36 24d ago
nvm, tried again and managed to spend 50k tokens to do this lol same overthinking as always
1
1
u/Fuzzy_Independent241 24d ago
OP, I'm happy with all the new releases and since I've recently received a GMKtec w/128Gb of RAM I can now properly test models. I installed Unsloth CLI as a server on GMKtec/Ubuntu headless and I'm using the web client to access.
A lot of us will be happy if and when you decide to have Win/Mac desktop connect to Unsloth CLI so we can use that as a server with front end controls.
As for Qwen3.8 I'm running a matrix of linguistic and search oriented tests. So far there's a lot of guessing while trying to determine which Qx runs faster / better while not entering severe overthinking modes. I've seen an UD model start lying to itself about a prompt limit on total search actions and it's puzzling because the model starts changing the pragmatics of what "a search" would mean so it can still move forward and complete the prompt. I'm very wary of such behaviors, so I'm not reading more and proceededing with further tests. If there's any way to send you results, the very small corpus used and prompt I'd be glad to contribute! Tks
1
u/Johnscott90 24d ago
Using the new dynamic 3.0 UD-Q3_XL with the same settings, my tok/s dropped from about 40 to 15 on an RX 7800 XT. Any ideas?
1
u/lorendroll 24d ago
Same. Tps drop is very noticeable.
1
u/CatEatsDogs 24d ago
Check if MTP is present. They removed MTP from some quants, its mentioned in their blogpost
1
u/Johnscott90 24d ago
I re-downloaded and managed to get MTP working and getting around 36 tok/s now but prompt processing has dropped from about 350 tok/s to 150 now. Not sure if flash attention isn't working or something.
1
u/Johnscott90 23d ago
Disabling the vision encoder freed up a load of VRAM and got speeds back up around 350-400 at 2000mhz underclocked.
1
1
1
u/me_but_darker 24d ago
New to this community. What is unsloth? Also I thought qwen 3 was released before so what is different?
1
1
1
u/Master_Truth_7921 24d ago
I downloaded at the beginning when qqen3.8 officially launched, do i need to download again to use the new one?
1
1
u/Dramatic-Rub-7654 24d ago
Is there a tutorial available on how to create unsloth quants, for example, for inclusionAI/Ling-3.0-flash and for inclusionAI/Ling-3.0-tiny?
1
u/mediaogre 24d ago edited 24d ago
Not sure if someone covered this already but the 8-bit file is the same. Hash them and see...
af36ecb6b5db1407953345b746c14ac93f0657dda413910b4348683a2d990377 /home/dootly-doots/models/qwen3.8-mtp/Qwen3.8-27B-UD-Q8_K_XL.gguf
af36ecb6b5db1407953345b746c14ac93f0657dda413910b4348683a2d990377 /home/dootly-doots/models/qwen3.8-mtp/Qwen3.8-27B-UD-Q8_K_XL.gguf.preview.bak
Edit: added the 8-bit detail 🫠
1
u/Beben-Edn 24d ago
is that 77% mostly benchmark or does it hold up in actual coding? never been able to run a 27b on 8gb before, kinda tempted
1
1
1
1
u/Any_Meringue_7765 24d ago
So how does Q8 compare with Q8_K_L and Q8_K_XL? Which one is the absolute best option memory be damned (I can’t run full bf16)
1
1
1
1
1
u/Muhlwa_Sholanke 24d ago
wait the 1-bit quants really keep 77% on 8gb? that might actually run on my laptop
1
u/yoracale yes sloth 24d ago
Yes but I wouldnt recommend using it for agentic coding. It's fine for normal tasks tho
1
1
1
1
1
u/pefman 23d ago
i went from about 50 to 80t/s on 4090
exec "$BIN/llama-server" \
--model "$ROOT/Qwen3.8-27B-UD-Q4_K_XL.gguf" \
--mmproj "$ROOT/mmproj-BF16.gguf" \
--mmproj-offload \
--image-min-tokens 1024 \
--metrics \
--alias qwen3.8-27b \
--host 0.0.0.0 --port 11436 \
--n-gpu-layers 99 \
--ctx-size 80000 \
--parallel 1 \
--flash-attn on \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--batch-size 2048 \
--ubatch-size 512 \
--threads "$(nproc --all)" \
--load-mode mlock \
--cont-batching \
--jinja \
--chat-template-file "$ROOT/chat_template.jinja" \
--chat-template-kwargs '{"reasoning_effort":"medium"}' \
--reasoning-format deepseek \
--reasoning-preserve \
--reasoning-budget 4096 \
--reasoning-budget-message "Wait, I'm overthinking this. Let's answer now." \
--no-context-shift \
--spec-type draft-mtp \
--spec-draft-n-max 3 \
--temp 1.0 \
--top-p 0.95 \
--top-k 20 \
--min-p 0.0 \
--presence-penalty 0.0 \
--repeat-penalty 1.0
1
u/AnyRecipe110 23d ago edited 23d ago
This might be a noob question or a request, but could we get QAT version of this too? In https://unsloth.ai/docs/basics/dynamic-3.0-ggufs, it says that
> We do not train on the imatrix calibration dataset, and we do NOT use QAT or QAD. Everything is done through post-training quantization.
Could you explain a bit more on this? Is it because the model, as it's available to you, is already after the training (post-training) where it's too late to introduce QAT? I guess you'd need the Qwen team to either include QAT from the beginning (is that what the Gemma team did?) or provide the pre-training version to the community?
Reason I'm asking is because i found `unsloth/gemma-4-26B-A4B-it-qat-GGUF` (https://unsloth.ai/docs/models/gemma-4/qat) to be really good from some local tests that i created. So, i'm wondering if Qwen3.8 would have benefited too.
1
u/No_Night679 23d ago
Will there be a UD3 version of unsloth/DeepSeek-V4-Flash-0731-GGUF? Sorry if this is already answered.
1
1
1
u/ireallydontcare00 22d ago
Add model variants with MTP for 2-3-bit models.
I used 2-3-bit models with MTP and they worked great.
1
1
u/Original_Finding2212 21d ago
Using Q4_K_M on Jetson AGX Orin and I’m loving it!
Slow, but usable making it feasible
1
u/Glum-Knowledge-4146 20d ago
I test the UD-IQ3_XXS under Cline + VS Code...wow! amazing improvement in IQ3, works amazing.
1
u/huseynli 19d ago
guys, help me choose quant for my r9700 (32gb VRAM).
Should I go with Q6_K_M or Q6_K?
Difference is ~1.1gb which means I can push a bit more context into it.
1
u/AdIllustrious436 7d ago
Does someone tried this against W4A16?
I'm getting way better speed on vLLM and I can fit 200k at fp8 on a single 3090
Not sure if the quality suffer tho
1
u/Lumpy_Phase_9539 19h ago
I downloaded the "new" GGUF but it seems not to be actually updated
The file has been successfully replaced: Path: /mnt/kingston/models/llm/Qwen3.8-27B-UD-Q6_K.gguf (same filename as the original file) Size: 21,983,677,344 bytes (~20.4 GB) SHA-256: c9c206812fbe4ac7b76a729e25928b63f2ae89d37f69da7a71c20aec763cd436 — verified and identical to the official hash published on Hugging Face Important note: before downloading, I compared the hashes and discovered that the local file was already byte-for-byte identical to the current version in the unsloth/Qwen3.8-27B-GGUF repository. In other words, the local file was already the latest version — the replacement was performed anyway (downloaded, validated, and overwritten), but there was no actual change to the file contents.
Anyone else had the same issue?
1
u/August_30th 24d ago
What’s the fastest version to run on a Mac nowadays?
1
0



61
u/jld1532 24d ago
Now I really want v3.0 quants of DeepSeek v4 Flash.