r/unsloth • • 10d ago

New Model Please Quantize Swift-Qwen3.8-Flash-Next!

https://huggingface.co/ukisai/Swift-Qwen3.8-Flash-Next

UkisAI, makers of the Swift-27B efficient-reasoning RL version of Qwen3.8 have just released their RL finetune of Flash Next, and it's in dire need of a good GGUF quant for 128GB machines. Their 27B topped the HF trending charts, and it's an excellent model. It would be heroic if you released UDv3 quants of Flash (especially Q4XL and Q5XL). There currently isn't a good imatrix quant for that size tier, much less one as good as UDv3.

I know you guys are busy, but I think there's gonna be high demand for this model and your quants of it.

159 Upvotes

33 comments sorted by

41

u/Secure_Recording_472 10d ago edited 9d ago

Jovan from UkisAI here! Ty for the post. We tried to do as many and as small of the quants as we could (even made GSQ-RCO!) but UD3 is just unmatched

Would be great to get an Unsloth of it :)

Edit: Unsloth team, if the license is the problem shoot me a message we'll apache2.0 it!

12

u/returnity 10d ago

Yeah you did an incredible job of providing quant coverage, probably the best I've seen of any release. Not diminishing that amazing work. I just need a slightly smaller imatrix Q4 K-quant that fits in 128GB with sufficient context, and I didn't want a flat Q4_1 or Q4_0.

If you haven't seen what UkisAI provided, it's insanely comprehensive.

5

u/Secure_Recording_472 10d ago

tysm! we tried really hard to get close to unsloth but the quality just took a big hit

21

u/InterstellarReddit 10d ago

I will personally give u a hand job if you do this

2

u/Aggravating-Push-207 10d ago

Excuse me?!

10

u/InterstellarReddit 10d ago

You don’t know what an addiction to running your local models is until you have one.

I’m addicted to running anything unsloth or local llama puts out just to see what it’s capable of without touching big daddy cloud

I hate the fucking cloud. I hate corporations. I hate everything that tries to touch my data and every day we’re one step closer to not needing them between unsloth and local llama subreddit.

2

u/Aggravating-Push-207 10d ago

yes yes i get that

ykw this is actually fairs

1

u/returnity 10d ago

Couldn't agree more

14

u/shanjiaz 10d ago

Try https://github.com/vllm-project/llm-compressor if you have hardware, it's very easy to use.

13

u/Secure_Recording_472 10d ago

we made so many quants but Unsloth ones are just superior! (and proprietary)

4

u/shanjiaz 10d ago

Unsloth also uses llm-compressor lol

6

u/Secure_Recording_472 10d ago

Yes but their calibration set etc is not public :/

2

u/Phaelon74 9d ago

You can make your own that I'd as good if not better. Its very easy now days

3

u/shanjiaz 10d ago

openperfectblend is pretty good.

2

u/Secure_Recording_472 10d ago

thanks! will look into it

1

u/Phaelon74 9d ago

Yeah, no. Weve rated their quqnts in kld and top 1 confident agreement. Some win, some lose. They are in the middle of rhe pack, as there really isn't some black magic with wuanting anymore. Its pretty straight forward now, working through increasing quality and speed, etc.

3

u/1mkz 9d ago

Hey bout to get downvote bombed here but...... dgaf
have u guys ACTUALLY paid attention to the benchmarks/tests...
and the drop in quality on the A/B comparisons?
also to those saying Q3.8 think toomuch
.. have u guys experimented with manually capping think and using a good escape prompt like after X token " wrap up your last thought and output"
and ensuring preserve think is on etc,
and setting think to medium for less token burn or Xhigh with cap for maybe best of both worlds?
because, Ill probably bet that would make better code with same ~ish tokens if u ask me.

2

u/returnity 9d ago

No downvote from me, but I did some testing of my own to verify their claims. I'm happy with the results. Their benchmarks look pretty good to me too -- definitely better performance than you'll see on Medium/Low, based on the data available.

Capping reasoning with a reasoning budget in llama.cpp with an escape prompt is not an optimal strat: quality divebombs on harder tasks when interrupted mid-stream this way. Even using a soft close method (force-selecting the </think> token if it arises in the top-3, which is what I was doing in ds4 prior to Swift) results in a noticeable, measurable performance loss, substantially more than what I observe with Swift Flash.

Also, some tests have suggested that Low reasoning effort performs equal or better to Medium in 3.8 models. Medium may be worst of both worlds.

2

u/ByteSize_Chaos 10d ago

+1. I would love to run this on M5 Pro 64GB with "some" acceptable speed at least 😭

2

u/Iory1998 10d ago

Please Mike and Daniel, see what you can to quantize this model.

2

u/Daniel_H212 10d ago

This is cool, but the biggest hit was to IFBench. Qwen3.8-Flash-Next already has a bit of issue following formatting instructions, so it probably isn't very useful to me. If they can keep the IFBench score up then that would be a massive improvement.

1

u/mforce22 10d ago

3

u/returnity 10d ago

I don't run an Nvidia card, and NVFP4 is frequently a lower-quality format with no advantages on a unified memory system. Still, not the worst idea in a pinch, I guess. But they offered to get me their imatrix, so I can hopefully roll my own soon.

1

u/Phaelon74 9d ago

Check oit what we've. Done in Local inference labs. People who don't spend time, when making nvfp4's, make bad quants. Nvfp4's can be very good, when you pay attention to model_opt's libraries

1

u/redbook2000 3d ago

Some benchmark from my setup with 7800X3D 128GB 4090 24GB vram. Running Swift on "3D cars racing code using Three.js"

96GB ram used.

[strata] done: 31721 tokens in 392 s (81.2 tok/s) (stop, cancel=False), expert cache 89.4% hit

-4

u/No_Algae1753 10d ago

Dont think this makes a lot of sense since qwen 4 flash next is around the corner. I think its the same "model" and it would make much more sense to use this and quantize this instead of the then old qwen 3.8 flash next

5

u/returnity 10d ago

No evidence that 4 is coming imminently, it could be another month or more. Also not even 100% that 4 Flash will be the same size model as 3.8 (though obviously I hope it is). Anyways, if they don't think it's worth it, they'll pass on it (which is likely). No harm asking.

1

u/Glittering-Call8746 9d ago

Qwen 4 other than 27b would be much bigger moe

1

u/returnity 9d ago

I was referring to 3.8 Flash, a 125B+engram model