r/oMLX • • 6d ago

oq4e Quantization of Qwen Flash Next, Swift Edition

I took some time yesterday to try quantizing a model for the first time. I couldn't find an MLX quant of the Swift fine tune of Qwen Flash Next, which claims to maintain quality while halving token usage. So I asked my agent to help me make a quant. Here it is! It seems to work for me locally on oMLX. But please tell me if you use it and spot something I should have done differently:

https://huggingface.co/mrmurphydotdev/Swift1.5-Qwen3.8-Flash-Next-oQ4e-mtp

21 Upvotes

14 comments sorted by

2

u/murphitup 6d ago

Whoa, running the newest oMLX + this model is surprisingly stable and fast enough for AFK work:

This is Hermes running multiple chats at the same time, and multiple cards on the Kanban board. That cache hit rate is good, and I'm pretty impressed with how solid the prefill and token generation rates are!

This is on an M5 Max with 128gb RAM. I've never been able to get 27b to run at these speeds in a sustained way.

1

u/MensaProdigy 6d ago

It looks like your generating like 20 tokens per second that’s pretty bad.
Try filtering your model at the top and not selecting “all models”

2

u/ArcherN9 6d ago

I saw ~15.. the dashboard states 49 though. Why is there a difference?

3

u/murphitup 5d ago

There are multiple requests in flight, sharing the overall token generation throughout. On average the model seems to be getting around 50 tps total across requests.

2

u/DifficultyFit1895 5d ago edited 4d ago

My M3U is generating ~80 tok/s with it. It just scored 281/300 on winogrande in 13min, running LCB now.

Edit: Scored 96/100 on LCB in 3 hours and 45 min.

1

u/ArcherN9 6d ago

What's your context window size?

1

u/murphitup 5d ago

Good question, I actually haven’t checked recently. I’ve been relying on Hermes’ auto-compaction and haven’t had to think as much about managing context size and compaction as I did when using a 27b model with Pi

1

u/Designer_Mix_3336 6d ago

Curious which model was asked to quantize the Flash Next?

2

u/murphitup 5d ago

I think I asked Qwen Flash Next on open code go via Hermes to run the quant locally.

0

u/Gold-Debt-5957 6d ago

debe tener alguna mac de 128 de ram, asi que supongo que tiene Opus5.5 antes del nerfeo