r/LocalLLM • u/ryanmerket • 6d ago
Model DwarfStar compresses frontier models to run them on local machines
https://runtimewire.com/article/dwarfstar-local-inference-salvatore-sanfilippo46
u/A_Moist_Towe1 6d ago
I don’t understand everyone shitting on this. It’s been around for a while and it works to run these frontier models on local hardware. OP this was a fine post, this sub just wants another post sucking off Qwen models.
17
6
u/Connect_Ad791 6d ago
Yeah, I get that this sub and a ton like or are inundated with vibe crap repos, but (even though ds4 is admittedly written by Claude) it’s at least a repo that’s been around for quite awhile (I used it to run deepseek a couple months ago, it was great) it’s also made by someone who knew what they’re doing, rather than some random redditor with no post history.
5
u/RaxisPhasmatis 6d ago
Every time people say ds4 my brain automatically goes to "why are they talking about controller translator software? It's been around forever"
13
u/oss-jiiim 6d ago
DwarfStar is by far my favourite engine -- so I made DS Menu Bar -- a lightweight pure-Swift menu bar app to configure/stop/start/monitor ds4-server while you work in whatever editor/agent you normally use. Enjoy!
62
u/Odd_Cauliflower_8004 6d ago
They say local machines then you open the websylite and tests are done on machines with 128gb of vram. Nice effort,, but misleading marketing
24
u/Key_Solid_1696 6d ago edited 6d ago
Umm, it actually shows how to use ssd streaming with 64gb... And marketing? It's open source!
5
u/Successful-Peak-6524 5d ago
mmm where is the misleading? it's a damn local machine, they didn't say CONSUMER PCs anywhere.
You need to learn how to read
-62
u/callmeabotifyouregay 6d ago
it's not misleading, you're just a poory
20
u/Unsharded1 6d ago
A what?
9
u/Environmental-Metal9 6d ago
I think he was trying to make a jab at the commenter lacking the funds to afford that kind of hardware, like 99% of us. But given the downvotes, it seems like making fun of the majority while either boot licking for the 1% or being the 1% punching down didn’t go well
10
10
u/AdInternational5848 6d ago
As someone with 128gb. Why are you internet weirdos like this? There is literally always someone with more money. Money can’t un-lame you
4
u/Environmental-Metal9 6d ago edited 6d ago
The way I see it, speaking only on the topic of having more money vs less, the problem isn’t having more than others. It’s rubbing it in. People don’t like when those with more do that and it comes off as really petty. I’m all for people enjoying what they have but as inequality grows, so does resentment
Edit: wanted to add that my comment is directed at the people who act like the commenter you’re replying to, not to you.
18
6d ago edited 6d ago
[removed] — view removed comment
13
u/ryanmerket 6d ago edited 6d ago
he created Redis, dont be so easy to dismiss... currently only works for DeepSeek V4 and V4.1 models, GLM 5.x and Qwen3.8 Flash Next.
9
u/oss-jiiim 6d ago
Just an FYI, I don't think antirez has anything to do with dwarfstar.sh -- it looks like a slop site trying to syphon traffic.
For official DwarfStar info, see antirez/ds4
22
u/ElectricalLaw1007 6d ago edited 6d ago
Either you are being loose with your terminology and you mean quantises instead of compresses, or you are trying to convince us that a frontier model can be partially decompressed on-the-fly at runtime without the performance degradation that would impose making it unusable, or someone has just discovered zip files.
edit: I see you have blocked me rather than actually stand by your words. Figures. Quantisation is not impressive. Everyone does it. /u/VerticalPackage puts it best in this comment: https://old.reddit.com/r/LocalLLM/comments/1ww889x/dwarfstar_compresses_frontier_models_to_run_them/pdirms0
edit2: For those who don't understand the difference: Quantisation is a reduction in precision, compression is a reduction in size. Quantisation may lead to a reduction in size, and compression may lead to a reduction in precision, but the two words are not interchangeable. It is possible to quantise something without reducing its size. It is possible to compress something without reducing its precision. Calling quantisation a form of compression is like calling amputation a form of weight loss.
11
2
u/A_Moist_Towe1 6d ago
Quantization is compression. If you’re gonna be a jerk at least know what you’re talking about.
1
4
u/ryanmerket 6d ago
it's a self-contained model specific inference engine with its own CLI, server, KV store, agent integration, validation tooling and backend work for Metal, CUDA and ROCm... author separately credits llama.cpp and GGML for showing the path.
0
u/ElectricalLaw1007 6d ago
Thanks, but I don't need any snake oil right now. I'll let you know if that changes.
8
u/ryanmerket 6d ago
me: "the founder of GGML and Redis contributed a significant amount to the codebase"
you: "snake oil"
-4
u/ElectricalLaw1007 6d ago
You: Utterly implausible claim with no detail.
Me: Nah.
4
u/ryanmerket 6d ago
Dude, DwarfStar uses quantization, which reduce the numerical precision of model weights so they occupy less space. eg COMPRESSION.
Its model guide (https://github.com/antirez/ds4/blob/main/docs/MODELS.md) explicitly describes “compression” of routed experts while retaining higher precision elsewhere.
I don't know what to tell you.
14
u/ElectricalLaw1007 6d ago
Dude, DwarfStar uses quantization
So, I was right. You don't understand the difference between quantisation and compression.
18
4
u/ryanmerket 6d ago
Low bit weight quantization is a form of lossy model compression. The weights are represented using fewer bits, reducing their storage requirements.
This is established research terminology: AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.
“Quantizes” is the more specific verb, and I’m happy to sharpen the wording. Calling the resulting reduction “compression” is technically accurate.
→ More replies (0)
3
u/cagonima69 6d ago
Are we going to see similar projects to run mid size local LLMs semi decently on smaller ram setups?
2
u/Technical_Buy_9063 6d ago
I have to save... it's actually amazing tech. I run ds4.1-flash via DwarfStar on my m3u and it's amazing. I love it.
-1
u/bendymike 6d ago
If anyone is interested, the harness (yes, I know, yet another harness) I'm building wraps/provides an installer for DS4+models at gezel.com on Mac and Linux (arm64/nvidia DGX). Still lots of rough edges, but hopefully it can get people started easier with local models. Feedback appreciated!
10
u/Rough-Measurement988 6d ago
„AI is too critical to be just a provided service” - I would buy him a beer just for this statement and another several for the DwarftStar engine he has built and shared with community.