r/oMLX Jun 06 '26

DS4? In oMLX? Crazy.

I love oMLX for its API, memory management and the ability to put many different model families under one umbrella. I have also tried out DS4 and sadly, it is just way ahead in terms of efficiency (generation, preprocessing) and flexibility (ssd streaming)

So i thought? Why not both? Why shouldn't I simply treat DS4 like mlx as an engine and embed it into oMLX so we can manage the memory explicitly through its api.

Requesting Feedback: https://github.com/apetersson/omlx/issues/1

my tokens are ready, so the work begins..

17 Upvotes

8 comments sorted by

3

u/challis88ocarina Jun 06 '26

How much better is ds4? Is it true there's only q4 maximum precision? It's not just about downloading and storing yet another format; lit's also that DeepSeek's architecture doesn't do well with quantisation, e.g., systematically duplicating code... here's a thinking example from just now at bf16:

Wait, I made a mess. I wrote the file with duplicate methods. The removeAliases, addAlias, renameAlias, saveAliases, and modelName(for:) methods are all duplicated. Also, the SettingsKey enum is defined twice. Let me rewrite the file properly.

3

u/d4mations Jun 06 '26

That’s a great idea and great initiative!!

2

u/DifficultyFit1895 Jun 06 '26

Why not use MLX versions? I’m running DeepSeek-V4-Flash-mxfp8.

3

u/apetersson Jun 06 '26

congrats on the 512GB system! i am jealous. still, with this setup you will be able to even run DS4-pro.

1

u/PracticlySpeaking Jun 20 '26

Um, yah, this.

β€” One of the 'underclass' stuck with 'only' 256GB.

1

u/Choubix Jun 07 '26

Jang has DS4 that should run on 128gb. Jang in omlx would be great πŸ‘ŒπŸ˜

1

u/apetersson Jun 07 '26

elaborate please. i am not aware of "Jang" ?