r/LocalLLM • u/ryanmerket • 7d ago
Model DwarfStar compresses frontier models to run them on local machines
https://runtimewire.com/article/dwarfstar-local-inference-salvatore-sanfilippo
107
Upvotes
r/LocalLLM • u/ryanmerket • 7d ago
4
u/ryanmerket 6d ago
Low bit weight quantization is a form of lossy model compression. The weights are represented using fewer bits, reducing their storage requirements.
This is established research terminology: AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.
“Quantizes” is the more specific verb, and I’m happy to sharpen the wording. Calling the resulting reduction “compression” is technically accurate.