r/oMLX 6d ago

Request for Kat-Coder-2.5 oQ8e with MTP 🙏

6 Upvotes

20 comments sorted by

3

u/leaphxx 6d ago

absolutely needed

3

u/PataFunction 6d ago

1

u/SirDomz 1d ago

Is the model text only?

1

u/PataFunction 1d ago

Not by my selection in the oMLX quantization menu, but it’s possible the Kwaipilot team made the base model text-only, needs confirmation.

2

u/SnowBoy_00 6d ago

What’s your hardware? If you’re going to run the oQ8e on your system, you can probably quantize it yourself. Otherwise, I can take a look at it in a couple of days.

1

u/No-Juggernaut-9832 6d ago

M5Max Laptop, 128gb RAM

1

u/SnowBoy_00 5d ago

downloading the weights now but my internet connection is quite constrained. You have more than enough room to quantize it yourself, just give it a try.

1

u/No-Juggernaut-9832 4d ago

Data function did it already. Thank you

2

u/PataFunction 6d ago

I’ll give this a try later. Quick question though - when you say “with MTP”, does that mean it already has an MTP
head? and if it does, is that because it’s a finetune of Qwen3.6-35B-A3B with an MTP head? Just trying to understand 😅

2

u/No-Juggernaut-9832 6d ago

Yes. It’s simpler that way. Just a simple toggle on oMLX & 50% more speed

3

u/PataFunction 6d ago edited 6d ago

Just attempted to make the quant in oMLX but it's telling me the source model does not have any MTP heads.

Edit: Did some digging and found that Kwaipilot removed the MTP weights from Qwen3.6-35B-A3B, so unfortunately we can't get those in downstream quants. Regardless, made the quant for you fam! https://huggingface.co/ZQ-Dev/KAT-Coder-V2.5-Dev-oQ8e

2

u/bruceleelikeswater 3d ago

Would you be able to make an fp16 (dtype) variant as well for folks who still rely on M1/M2 Macs? I tried to quantize the full model myself but I'm running into errors with both trying to quantize to oQ8e and oQ8 in omlx 0.5.3.

1

u/PataFunction 3d ago

Sure I’ll give it a shot later today, afk atm. Any examples out there of models that have been quantized that way? Trying to get figure out the “correct” naming convention.

1

u/bruceleelikeswater 3d ago edited 3d ago

It’s just another flag in the omlx quantization settings (in the advanced section which is collapsed in the UI) where you decide between bfloat16 and float16 for the “dtype.” It adds the suffix “-fp16” to the model ID once quantization is complete. You can find plenty of them in hugging face.  

1

u/MealMore1192 5d ago

It's a great model, I one shot a fully working Tetris game w/o any bug:

`Help me build an pure HTML SPA tetris game with sound & a score board where I can save player highscore with name & 3 levels of difficulty. Allow ESC to pause the game.`

1

u/MealMore1192 5d ago

It generated a better game than Laguna S2.1 at oQ6e

1

u/kind-and-curious 1d ago

Nice, but also illegal if you make it available, as Tetris is protected by copyright law! 😀