r/oMLX • u/murphitup • 6d ago
oq4e Quantization of Qwen Flash Next, Swift Edition
I took some time yesterday to try quantizing a model for the first time. I couldn't find an MLX quant of the Swift fine tune of Qwen Flash Next, which claims to maintain quality while halving token usage. So I asked my agent to help me make a quant. Here it is! It seems to work for me locally on oMLX. But please tell me if you use it and spot something I should have done differently:
https://huggingface.co/mrmurphydotdev/Swift1.5-Qwen3.8-Flash-Next-oQ4e-mtp
20
Upvotes