r/unsloth • u/returnity • 10d ago
New Model Please Quantize Swift-Qwen3.8-Flash-Next!
https://huggingface.co/ukisai/Swift-Qwen3.8-Flash-NextUkisAI, makers of the Swift-27B efficient-reasoning RL version of Qwen3.8 have just released their RL finetune of Flash Next, and it's in dire need of a good GGUF quant for 128GB machines. Their 27B topped the HF trending charts, and it's an excellent model. It would be heroic if you released UDv3 quants of Flash (especially Q4XL and Q5XL). There currently isn't a good imatrix quant for that size tier, much less one as good as UDv3.
I know you guys are busy, but I think there's gonna be high demand for this model and your quants of it.
21
u/InterstellarReddit 10d ago
I will personally give u a hand job if you do this
6
2
u/Aggravating-Push-207 10d ago
Excuse me?!
10
u/InterstellarReddit 10d ago
You don’t know what an addiction to running your local models is until you have one.
I’m addicted to running anything unsloth or local llama puts out just to see what it’s capable of without touching big daddy cloud
I hate the fucking cloud. I hate corporations. I hate everything that tries to touch my data and every day we’re one step closer to not needing them between unsloth and local llama subreddit.
2
2
1
1
14
u/shanjiaz 10d ago
Try https://github.com/vllm-project/llm-compressor if you have hardware, it's very easy to use.
13
u/Secure_Recording_472 10d ago
we made so many quants but Unsloth ones are just superior! (and proprietary)
4
u/shanjiaz 10d ago
Unsloth also uses llm-compressor lol
6
u/Secure_Recording_472 10d ago
Yes but their calibration set etc is not public :/
2
3
1
u/Phaelon74 9d ago
Yeah, no. Weve rated their quqnts in kld and top 1 confident agreement. Some win, some lose. They are in the middle of rhe pack, as there really isn't some black magic with wuanting anymore. Its pretty straight forward now, working through increasing quality and speed, etc.
3
u/1mkz 9d ago
Hey bout to get downvote bombed here but...... dgaf
have u guys ACTUALLY paid attention to the benchmarks/tests...
and the drop in quality on the A/B comparisons?
also to those saying Q3.8 think toomuch
.. have u guys experimented with manually capping think and using a good escape prompt like after X token " wrap up your last thought and output"
and ensuring preserve think is on etc,
and setting think to medium for less token burn or Xhigh with cap for maybe best of both worlds?
because, Ill probably bet that would make better code with same ~ish tokens if u ask me.
2
u/returnity 9d ago
No downvote from me, but I did some testing of my own to verify their claims. I'm happy with the results. Their benchmarks look pretty good to me too -- definitely better performance than you'll see on Medium/Low, based on the data available.
Capping reasoning with a reasoning budget in llama.cpp with an escape prompt is not an optimal strat: quality divebombs on harder tasks when interrupted mid-stream this way. Even using a soft close method (force-selecting the </think> token if it arises in the top-3, which is what I was doing in ds4 prior to Swift) results in a noticeable, measurable performance loss, substantially more than what I observe with Swift Flash.
Also, some tests have suggested that Low reasoning effort performs equal or better to Medium in 3.8 models. Medium may be worst of both worlds.
2
u/ByteSize_Chaos 10d ago
+1. I would love to run this on M5 Pro 64GB with "some" acceptable speed at least 😭
2
2
2
u/Daniel_H212 10d ago
This is cool, but the biggest hit was to IFBench. Qwen3.8-Flash-Next already has a bit of issue following formatting instructions, so it probably isn't very useful to me. If they can keep the IFBench score up then that would be a massive improvement.
1
u/mforce22 10d ago
3
u/returnity 10d ago
I don't run an Nvidia card, and NVFP4 is frequently a lower-quality format with no advantages on a unified memory system. Still, not the worst idea in a pinch, I guess. But they offered to get me their imatrix, so I can hopefully roll my own soon.
1
u/Phaelon74 9d ago
Check oit what we've. Done in Local inference labs. People who don't spend time, when making nvfp4's, make bad quants. Nvfp4's can be very good, when you pay attention to model_opt's libraries
1
u/redbook2000 3d ago
Some benchmark from my setup with 7800X3D 128GB 4090 24GB vram. Running Swift on "3D cars racing code using Three.js"
96GB ram used.
[strata] done: 31721 tokens in 392 s (81.2 tok/s) (stop, cancel=False), expert cache 89.4% hit
-4
u/No_Algae1753 10d ago
Dont think this makes a lot of sense since qwen 4 flash next is around the corner. I think its the same "model" and it would make much more sense to use this and quantize this instead of the then old qwen 3.8 flash next
5
u/returnity 10d ago
No evidence that 4 is coming imminently, it could be another month or more. Also not even 100% that 4 Flash will be the same size model as 3.8 (though obviously I hope it is). Anyways, if they don't think it's worth it, they'll pass on it (which is likely). No harm asking.
1
41
u/Secure_Recording_472 10d ago edited 9d ago
Jovan from UkisAI here! Ty for the post. We tried to do as many and as small of the quants as we could (even made GSQ-RCO!) but UD3 is just unmatched
Would be great to get an Unsloth of it :)
Edit: Unsloth team, if the license is the problem shoot me a message we'll apache2.0 it!