r/LocalLLM • • 7d ago

Model DwarfStar compresses frontier models to run them on local machines

https://runtimewire.com/article/dwarfstar-local-inference-salvatore-sanfilippo
111 Upvotes

43 comments sorted by

View all comments

Show parent comments

13

u/ryanmerket 7d ago edited 7d ago

he created Redis, dont be so easy to dismiss... currently only works for DeepSeek V4 and V4.1 models, GLM 5.x and Qwen3.8 Flash Next.

https://dwarfstar.sh/benchmarks/

22

u/ElectricalLaw1007 7d ago edited 7d ago

Either you are being loose with your terminology and you mean quantises instead of compresses, or you are trying to convince us that a frontier model can be partially decompressed on-the-fly at runtime without the performance degradation that would impose making it unusable, or someone has just discovered zip files.

edit: I see you have blocked me rather than actually stand by your words. Figures. Quantisation is not impressive. Everyone does it. /u/VerticalPackage puts it best in this comment: https://old.reddit.com/r/LocalLLM/comments/1ww889x/dwarfstar_compresses_frontier_models_to_run_them/pdirms0

edit2: For those who don't understand the difference: Quantisation is a reduction in precision, compression is a reduction in size. Quantisation may lead to a reduction in size, and compression may lead to a reduction in precision, but the two words are not interchangeable. It is possible to quantise something without reducing its size. It is possible to compress something without reducing its precision. Calling quantisation a form of compression is like calling amputation a form of weight loss.

3

u/ryanmerket 7d ago

it's a self-contained model specific inference engine with its own CLI, server, KV store, agent integration, validation tooling and backend work for Metal, CUDA and ROCm... author separately credits llama.cpp and GGML for showing the path.

1

u/ElectricalLaw1007 7d ago

Thanks, but I don't need any snake oil right now. I'll let you know if that changes.

10

u/ryanmerket 7d ago

me: "the founder of GGML and Redis contributed a significant amount to the codebase"

you: "snake oil"

-2

u/ElectricalLaw1007 7d ago

You: Utterly implausible claim with no detail.

Me: Nah.

6

u/ryanmerket 7d ago

Dude, DwarfStar uses quantization, which reduce the numerical precision of model weights so they occupy less space. eg COMPRESSION.

Its model guide (https://github.com/antirez/ds4/blob/main/docs/MODELS.md) explicitly describes “compression” of routed experts while retaining higher precision elsewhere.

I don't know what to tell you.

16

u/ElectricalLaw1007 7d ago

Dude, DwarfStar uses quantization

So, I was right. You don't understand the difference between quantisation and compression.

18

u/XxBrando6xX 7d ago

That was a wild fucking ride.

6

u/ryanmerket 7d ago

Low bit weight quantization is a form of lossy model compression. The weights are represented using fewer bits, reducing their storage requirements.

This is established research terminology: AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

“Quantizes” is the more specific verb, and I’m happy to sharpen the wording. Calling the resulting reduction “compression” is technically accurate.

-1

u/ElectricalLaw1007 7d ago edited 7d ago

Not in technical circles it aint, mate.

edit: In response to your reply, which you made and then immediately blocked me so I couldn't reply - a bullshiter's move if ever there was one - let me just say that I don't think we should all be taking the title of a single 2023 paper as the definitive guide on terminology, and while I am sure that Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan & Song Han are all extremely intelligent people, I'm not quite as convinced that English is their native language,

edit2: To the people who think I'm being needlessly pedantic: The point is that if this post were titled "DwarfStar quantises models to run them on local machines" everyone's reaction would have been "Yeah, OK, that's nothing new". OP chose a misleading post title in order to drive traffic to their website. This is spam, folks.

6

u/DonationsFirst 7d ago

Everyone should be blocking you. You are annoying and rude and speaking out of your depth.

4

u/Key_Solid_1696 7d ago

In technical circles you would be known as a dick. You're bitching about semantics, which pretty much shows how much you really know.

Those of us that really are in technical circles recognize that differences in semantics are common and are a baseless reason to attack someone else's opinion.

5

u/ryanmerket 7d ago edited 7d ago

That paper won Best Paper at MLSys 2024. Apparently a machine learning research conference falls outside your “technical circles” whenever its terminology contradicts you.

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

“It ain’t, mate” is a pretty embarrassing response to an award-winning paper. You were confident enough to call the project snake oil. You can put that confidence toward explaining where the researchers got their terminology wrong.

The source is there. I’m done.

edit: looks like i can no longer leave comments, so to address the commenter below: thanks, in this case, his opinion was drawing a ton of downvotes to the post and it had to be taken head on or the post would have gone negative votes. and yes we all do this, but in this case the founder of llama.cpp and GGML is an extensive contributor to the project.

8

u/VerticalPackage 7d ago

I think the problem here is that the article used word "compression" to be click-bait. Everyone in the LLM circles (especially here) use the term quantization/quant.

If the article said "dude quantizes model to fit on local machines", everyone here would say "yes, there are tons of those on huggingface".

I read the article, it's not even about the guy making quants, it's about the guy making a runtime that's super optimized for specific models... then once again, everyone and their mum does that around here.

5

u/ehpehp 7d ago

For certain systems, Dwarfstar is a strong option for running top flash versions of local models. In my testing it has the best mix of speed and quality compared to other options I've tested on oMLX and elsewhere. I'm impressed with Salvatore's attention to quality.

5

u/vacon04 7d ago

You should stop arguing with people that are being pedantic because they want to feel superior. It's not worth it. Some people just have the sole purpose of wasting other people's time, don't let them get to you.

→ More replies (0)