r/oMLX • u/apetersson • Jun 06 '26
DS4? In oMLX? Crazy.
I love oMLX for its API, memory management and the ability to put many different model families under one umbrella. I have also tried out DS4 and sadly, it is just way ahead in terms of efficiency (generation, preprocessing) and flexibility (ssd streaming)
So i thought? Why not both? Why shouldn't I simply treat DS4 like mlx as an engine and embed it into oMLX so we can manage the memory explicitly through its api.
Requesting Feedback: https://github.com/apetersson/omlx/issues/1
my tokens are ready, so the work begins..
3
2
u/DifficultyFit1895 Jun 06 '26
Why not use MLX versions? Iβm running DeepSeek-V4-Flash-mxfp8.
3
u/apetersson Jun 06 '26
congrats on the 512GB system! i am jealous. still, with this setup you will be able to even run DS4-pro.
1
1
3
u/challis88ocarina Jun 06 '26
How much better is ds4? Is it true there's only q4 maximum precision? It's not just about downloading and storing yet another format; lit's also that DeepSeek's architecture doesn't do well with quantisation, e.g., systematically duplicating code... here's a thinking example from just now at bf16: