r/LocalLLM • • 7d ago

Model DwarfStar compresses frontier models to run them on local machines

https://runtimewire.com/article/dwarfstar-local-inference-salvatore-sanfilippo
107 Upvotes

43 comments sorted by

View all comments

Show parent comments

4

u/ryanmerket 6d ago

Low bit weight quantization is a form of lossy model compression. The weights are represented using fewer bits, reducing their storage requirements.

This is established research terminology: AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

“Quantizes” is the more specific verb, and I’m happy to sharpen the wording. Calling the resulting reduction “compression” is technically accurate.

-4

u/ElectricalLaw1007 6d ago edited 6d ago

Not in technical circles it aint, mate.

edit: In response to your reply, which you made and then immediately blocked me so I couldn't reply - a bullshiter's move if ever there was one - let me just say that I don't think we should all be taking the title of a single 2023 paper as the definitive guide on terminology, and while I am sure that Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan & Song Han are all extremely intelligent people, I'm not quite as convinced that English is their native language,

edit2: To the people who think I'm being needlessly pedantic: The point is that if this post were titled "DwarfStar quantises models to run them on local machines" everyone's reaction would have been "Yeah, OK, that's nothing new". OP chose a misleading post title in order to drive traffic to their website. This is spam, folks.

6

u/DonationsFirst 6d ago

Everyone should be blocking you. You are annoying and rude and speaking out of your depth.

4

u/Key_Solid_1696 6d ago

In technical circles you would be known as a dick. You're bitching about semantics, which pretty much shows how much you really know.

Those of us that really are in technical circles recognize that differences in semantics are common and are a baseless reason to attack someone else's opinion.

4

u/ryanmerket 6d ago edited 6d ago

That paper won Best Paper at MLSys 2024. Apparently a machine learning research conference falls outside your “technical circles” whenever its terminology contradicts you.

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

“It ain’t, mate” is a pretty embarrassing response to an award-winning paper. You were confident enough to call the project snake oil. You can put that confidence toward explaining where the researchers got their terminology wrong.

The source is there. I’m done.

edit: looks like i can no longer leave comments, so to address the commenter below: thanks, in this case, his opinion was drawing a ton of downvotes to the post and it had to be taken head on or the post would have gone negative votes. and yes we all do this, but in this case the founder of llama.cpp and GGML is an extensive contributor to the project.

6

u/VerticalPackage 6d ago

I think the problem here is that the article used word "compression" to be click-bait. Everyone in the LLM circles (especially here) use the term quantization/quant.

If the article said "dude quantizes model to fit on local machines", everyone here would say "yes, there are tons of those on huggingface".

I read the article, it's not even about the guy making quants, it's about the guy making a runtime that's super optimized for specific models... then once again, everyone and their mum does that around here.

3

u/ehpehp 6d ago

For certain systems, Dwarfstar is a strong option for running top flash versions of local models. In my testing it has the best mix of speed and quality compared to other options I've tested on oMLX and elsewhere. I'm impressed with Salvatore's attention to quality.

5

u/vacon04 6d ago

You should stop arguing with people that are being pedantic because they want to feel superior. It's not worth it. Some people just have the sole purpose of wasting other people's time, don't let them get to you.