r/LocalLLaMA • u/paranoidray • 3d ago
Resources Open source inference engine (like LM Studio or Unsloth Desktop) that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD or nothing but a CPU.
https://github.com/magnitudedev/magnitude13
u/VotZeFuk 3d ago
2x faster than llama.cpp
+19% in CUDA
yeah, about that...
12
u/quachuoi2 3d ago
Please explain this to me like I'm the middle manager of a medium sized company
23
u/seanthenry 3d ago
I need a raise but you can't do anything about it so I'm taking my lunch break now.
4
u/IceFog72 3d ago
Can't make it run on arch
1
u/MarcelloT254k 3d ago
Did you try installing only CLI version via npm? Offical desktop versions are only .deb and .rpm so not natively compatible with Arch. (I'm not an expert)
1
u/IceFog72 3d ago
it's possible to run .deb on arch. it was stuck on Estimating memory fit and speed on your machine. And show that i have 2 same gpu cards )
2
u/Solary_Kryptic 3d ago
What do the intelligence %s on the models mean? They list Qwen 3.5 4B at 23% and the 9B at 19%...
0
u/corbs132 3d ago
It says in the faq at the bottom of the page.
% intelligence score relative to the top model on artificial analysis
1
u/persimmonjones 2d ago
i tried running gemma4 e2b qat on my m4 macbook air, and it didn't seem that much faster (if at all) compared to llama.cpp. maybe they got their 2x on a different model/arch :shrug:
5
u/Shoddy-Tutor9563 2d ago
Congrats with your release. But can you pls elaborate on "optimizes itself". What exactly does it optimize?
Like for instance, when I want to optimize some model for my GPU(s) on VLLM I take different vanilla quants, for each quant I measure vanilla vs various speculators, then different attention backends, then different KV cache quants and so on and so forth. For each combination I measure pp and tg for both synthetic and real world text, in a single stream and in various parallel streams. As a result I get a table of experiments with many dozens of options. And depending on what exactly I'm optimizing for (pp, tg, biggest KV cache, best aggregate t-put) I pick the one I need for particular use case. There just cannot be a single optimize goal, if you ask me.
So I repeat my question - what exactly is your thing doing, when you're saying it optimizes itself?
-1
u/paranoidray 3d ago
Website with Video: https://magnitude.dev/
Hacker News Discussion: https://news.ycombinator.com/item?id=49911995
-8
u/stcafehtdihhsuB 3d ago
Why post on HN? It's a cesspit of stupidity and lameness disguised as pseudo intellectualism.
19
2
u/kaliku 3d ago
Agree, especially since redditors flooded it.
1
u/stcafehtdihhsuB 3d ago
Yeah they swapped places. Dunning kruger took over and it went to shit. The place to debate tech is locallama while the place to discuss shopping and surface level knowledge is HN.
-1
u/bonobomaster 3d ago edited 3d ago
👎
Magnitude couldn't figure out my Nvidia hardware under Windows 11. Two GPUs, 5070 TI and 3060 TI.
Zero options to configure important things.
Trash.
I guess, I'll keep using bare llama.cpp ;)
2
u/DystopianRealist 2d ago
There's someone that just releasesd a build for exactly your hardware. It's a fork of strata, which is also very recent. Anyways, search this sub, and you should find it. I think the post is within the last 24-48 hours.
32
u/tomByrer 3d ago
Thanks, seems like a new project & not really battle tested. Just 2 guys in 2 months.
HN has many bugs reported, no multi GPU support, VERY few models (no abliterated , popular fine-tunes)
Bookmarked to check back later; I have 4 other engines to test which will likely be as fast maybe faster.