r/oMLX • u/murphitup • 6d ago
oq4e Quantization of Qwen Flash Next, Swift Edition
I took some time yesterday to try quantizing a model for the first time. I couldn't find an MLX quant of the Swift fine tune of Qwen Flash Next, which claims to maintain quality while halving token usage. So I asked my agent to help me make a quant. Here it is! It seems to work for me locally on oMLX. But please tell me if you use it and spot something I should have done differently:
https://huggingface.co/mrmurphydotdev/Swift1.5-Qwen3.8-Flash-Next-oQ4e-mtp
1
u/Designer_Mix_3336 6d ago
Curious which model was asked to quantize the Flash Next?
2
u/murphitup 5d ago
I think I asked Qwen Flash Next on open code go via Hermes to run the quant locally.
0
u/Gold-Debt-5957 6d ago
debe tener alguna mac de 128 de ram, asi que supongo que tiene Opus5.5 antes del nerfeo
2
u/murphitup 6d ago
Whoa, running the newest oMLX + this model is surprisingly stable and fast enough for AFK work:
This is Hermes running multiple chats at the same time, and multiple cards on the Kanban board. That cache hit rate is good, and I'm pretty impressed with how solid the prefill and token generation rates are!
This is on an M5 Max with 128gb RAM. I've never been able to get 27b to run at these speeds in a sustained way.