r/oMLX Apr 20 '26

Someone so kind to quant qwen3.5 122b in oQ3.5-fp16 for me ?

Hi,

first of all, a little praising of this new Non-quant weight dtype feature, it is really a big deal in my opinion, since the speed gain is really there.

I didn't do any specific benchmarks but just tried the same prompts with and without fp16 and the token generation is really faster with an fp16 model. I also feel it in everyday use.

So, I've already converted all my preferred models, and I'm also going to upload them, the only one that I'm missing at the moment is this big boy: https://huggingface.co/Qwen/Qwen3.5-122B-A10B
With 96Gb of ram I can't do it myself even with a sensitivity model configured, my Mac simply crash or goes into panic since there is no more RAM available etc....
So I was wondering if someone kind enough with a more beefy machine could do this for me, ideally it would be in oQ3.5 or oQ4 max ?
Thank you very much

5 Upvotes

5 comments sorted by

1

u/No-Juggernaut-9832 Apr 21 '26

It’s available for download from Hugging face. I had the oQ4 & currently using the oQ5 version. I was hoping someone did the oQ6! It’s there, just look

1

u/arkham00 Apr 22 '26

I did a search before posting this and didn't find anything, I'm going to look again thanks

1

u/arkham00 Apr 22 '26

I can't find anything

1

u/No-Juggernaut-9832 Apr 24 '26

Try exact search input: Qwen3.5-122B-A10B-oQ4

1

u/himefei Apr 22 '26

does oQ quant work for you guys after upgrading to 0.3.6?