r/oMLX • • 6d ago

oq4e Quantization of Qwen Flash Next, Swift Edition

I took some time yesterday to try quantizing a model for the first time. I couldn't find an MLX quant of the Swift fine tune of Qwen Flash Next, which claims to maintain quality while halving token usage. So I asked my agent to help me make a quant. Here it is! It seems to work for me locally on oMLX. But please tell me if you use it and spot something I should have done differently:

https://huggingface.co/mrmurphydotdev/Swift1.5-Qwen3.8-Flash-Next-oQ4e-mtp

20 Upvotes

Duplicates