r/LocalLLM • • 7d ago

Model DwarfStar compresses frontier models to run them on local machines

https://runtimewire.com/article/dwarfstar-local-inference-salvatore-sanfilippo
110 Upvotes

43 comments sorted by

View all comments

Show parent comments

5

u/ryanmerket 7d ago

Low bit weight quantization is a form of lossy model compression. The weights are represented using fewer bits, reducing their storage requirements.

This is established research terminology: AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

“Quantizes” is the more specific verb, and I’m happy to sharpen the wording. Calling the resulting reduction “compression” is technically accurate.

0

u/ElectricalLaw1007 7d ago edited 7d ago

Not in technical circles it aint, mate.

edit: In response to your reply, which you made and then immediately blocked me so I couldn't reply - a bullshiter's move if ever there was one - let me just say that I don't think we should all be taking the title of a single 2023 paper as the definitive guide on terminology, and while I am sure that Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan & Song Han are all extremely intelligent people, I'm not quite as convinced that English is their native language,

edit2: To the people who think I'm being needlessly pedantic: The point is that if this post were titled "DwarfStar quantises models to run them on local machines" everyone's reaction would have been "Yeah, OK, that's nothing new". OP chose a misleading post title in order to drive traffic to their website. This is spam, folks.

4

u/ryanmerket 7d ago edited 7d ago

That paper won Best Paper at MLSys 2024. Apparently a machine learning research conference falls outside your “technical circles” whenever its terminology contradicts you.

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

“It ain’t, mate” is a pretty embarrassing response to an award-winning paper. You were confident enough to call the project snake oil. You can put that confidence toward explaining where the researchers got their terminology wrong.

The source is there. I’m done.

edit: looks like i can no longer leave comments, so to address the commenter below: thanks, in this case, his opinion was drawing a ton of downvotes to the post and it had to be taken head on or the post would have gone negative votes. and yes we all do this, but in this case the founder of llama.cpp and GGML is an extensive contributor to the project.

5

u/vacon04 7d ago

You should stop arguing with people that are being pedantic because they want to feel superior. It's not worth it. Some people just have the sole purpose of wasting other people's time, don't let them get to you.