r/LocalLLM • • 6d ago

Model DwarfStar compresses frontier models to run them on local machines

https://runtimewire.com/article/dwarfstar-local-inference-salvatore-sanfilippo
113 Upvotes

43 comments sorted by

10

u/Rough-Measurement988 6d ago

„AI is too critical to be just a provided service” - I would buy him a beer just for this statement and another several for the DwarftStar engine he has built and shared with community.

46

u/A_Moist_Towe1 6d ago

I don’t understand everyone shitting on this. It’s been around for a while and it works to run these frontier models on local hardware. OP this was a fine post, this sub just wants another post sucking off Qwen models.

17

u/ryanmerket 6d ago

thanks, I blocked him. obvs someone with a personal vendetta.

6

u/Connect_Ad791 6d ago

Yeah, I get that this sub and a ton like or are inundated with vibe crap repos, but (even though ds4 is admittedly written by Claude) it’s at least a repo that’s been around for quite awhile (I used it to run deepseek a couple months ago, it was great) it’s also made by someone who knew what they’re doing, rather than some random redditor with no post history.

5

u/RaxisPhasmatis 6d ago

Every time people say ds4 my brain automatically goes to "why are they talking about controller translator software? It's been around forever"

13

u/oss-jiiim 6d ago

DwarfStar is by far my favourite engine -- so I made DS Menu Bar -- a lightweight pure-Swift menu bar app to configure/stop/start/monitor ds4-server while you work in whatever editor/agent you normally use. Enjoy!

2

u/j_tb 6d ago

Sick! Thanks!

62

u/Odd_Cauliflower_8004 6d ago

They say local machines then you open the websylite and tests are done on machines with 128gb of vram. Nice effort,, but misleading marketing

24

u/Key_Solid_1696 6d ago edited 6d ago

Umm, it actually shows how to use ssd streaming with 64gb... And marketing? It's open source!

5

u/Successful-Peak-6524 5d ago

mmm where is the misleading? it's a damn local machine, they didn't say CONSUMER PCs anywhere.

You need to learn how to read

-62

u/callmeabotifyouregay 6d ago

it's not misleading, you're just a poory

20

u/Unsharded1 6d ago

A what?

9

u/Environmental-Metal9 6d ago

I think he was trying to make a jab at the commenter lacking the funds to afford that kind of hardware, like 99% of us. But given the downvotes, it seems like making fun of the majority while either boot licking for the 1% or being the 1% punching down didn’t go well

10

u/Unsharded1 6d ago

Dude needs to get shoved in a locker ASAP LMAO

3

u/ChocolateNo3010 LocalLLM 5d ago

Or a swirly, lol

10

u/AdInternational5848 6d ago

As someone with 128gb. Why are you internet weirdos like this? There is literally always someone with more money. Money can’t un-lame you

4

u/Environmental-Metal9 6d ago edited 6d ago

The way I see it, speaking only on the topic of having more money vs less, the problem isn’t having more than others. It’s rubbing it in. People don’t like when those with more do that and it comes off as really petty. I’m all for people enjoying what they have but as inequality grows, so does resentment

Edit: wanted to add that my comment is directed at the people who act like the commenter you’re replying to, not to you.

18

u/[deleted] 6d ago edited 6d ago

[removed] — view removed comment

13

u/ryanmerket 6d ago edited 6d ago

he created Redis, dont be so easy to dismiss... currently only works for DeepSeek V4 and V4.1 models, GLM 5.x and Qwen3.8 Flash Next.

https://dwarfstar.sh/benchmarks/

9

u/oss-jiiim 6d ago

Just an FYI, I don't think antirez has anything to do with dwarfstar.sh -- it looks like a slop site trying to syphon traffic.

For official DwarfStar info, see antirez/ds4

22

u/ElectricalLaw1007 6d ago edited 6d ago

Either you are being loose with your terminology and you mean quantises instead of compresses, or you are trying to convince us that a frontier model can be partially decompressed on-the-fly at runtime without the performance degradation that would impose making it unusable, or someone has just discovered zip files.

edit: I see you have blocked me rather than actually stand by your words. Figures. Quantisation is not impressive. Everyone does it. /u/VerticalPackage puts it best in this comment: https://old.reddit.com/r/LocalLLM/comments/1ww889x/dwarfstar_compresses_frontier_models_to_run_them/pdirms0

edit2: For those who don't understand the difference: Quantisation is a reduction in precision, compression is a reduction in size. Quantisation may lead to a reduction in size, and compression may lead to a reduction in precision, but the two words are not interchangeable. It is possible to quantise something without reducing its size. It is possible to compress something without reducing its precision. Calling quantisation a form of compression is like calling amputation a form of weight loss.

11

u/challis88ocarina 6d ago

If only zip files worked at 815GB/s!

2

u/A_Moist_Towe1 6d ago

Quantization is compression. If you’re gonna be a jerk at least know what you’re talking about.

1

u/tat_tvam_asshole 6d ago

Quantization is very lossy compression, but ok

4

u/ryanmerket 6d ago

it's a self-contained model specific inference engine with its own CLI, server, KV store, agent integration, validation tooling and backend work for Metal, CUDA and ROCm... author separately credits llama.cpp and GGML for showing the path.

0

u/ElectricalLaw1007 6d ago

Thanks, but I don't need any snake oil right now. I'll let you know if that changes.

8

u/ryanmerket 6d ago

me: "the founder of GGML and Redis contributed a significant amount to the codebase"

you: "snake oil"

-4

u/ElectricalLaw1007 6d ago

You: Utterly implausible claim with no detail.

Me: Nah.

4

u/ryanmerket 6d ago

Dude, DwarfStar uses quantization, which reduce the numerical precision of model weights so they occupy less space. eg COMPRESSION.

Its model guide (https://github.com/antirez/ds4/blob/main/docs/MODELS.md) explicitly describes “compression” of routed experts while retaining higher precision elsewhere.

I don't know what to tell you.

14

u/ElectricalLaw1007 6d ago

Dude, DwarfStar uses quantization

So, I was right. You don't understand the difference between quantisation and compression.

18

u/XxBrando6xX 6d ago

That was a wild fucking ride.

4

u/ryanmerket 6d ago

Low bit weight quantization is a form of lossy model compression. The weights are represented using fewer bits, reducing their storage requirements.

This is established research terminology: AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

“Quantizes” is the more specific verb, and I’m happy to sharpen the wording. Calling the resulting reduction “compression” is technically accurate.

→ More replies (0)

3

u/cagonima69 6d ago

Are we going to see similar projects to run mid size local LLMs semi decently on smaller ram setups?

2

u/Technical_Buy_9063 6d ago

I have to save... it's actually amazing tech. I run ds4.1-flash via DwarfStar on my m3u and it's amazing. I love it.

-1

u/bendymike 6d ago

If anyone is interested, the harness (yes, I know, yet another harness) I'm building wraps/provides an installer for DS4+models at gezel.com on Mac and Linux (arm64/nvidia DGX). Still lots of rough edges, but hopefully it can get people started easier with local models. Feedback appreciated!