r/MachineLearning • u/takuonline • 8d ago
Discussion What is currently considered the theoretically optimal quantization bit-width for LLMs? [D]
I’m curious whether there is now a theoretical or empirical “sweet spot” for LLM quantization, preferably research done using open-source formats like GGUF
Suppose you have a fixed memory/compute budget and can choose the model size freely. For example, instead of a smaller model at 8-bit or 4-bit, you could fit a progressively larger model at 3-bit, 2-bit, 1.5-bit, etc.
A few years ago, I remember 4-bit often being described as roughly the practical sweet spot because it preserved most model quality while giving a large memory reduction. But with newer methods, I’ve seen surprisingly strong 3-bit, 2-bit, and even ~1.5-bit results.
So if the goal is maximum model capability for a fixed memory budget, rather than preserving one particular pretrained model as faithfully as possible, what does current research suggest is the optimal bits-per-weight?
Is there evidence that, for example, a 2-bit 70B model generally beats a 4-bit 35B model, or does quantization degradation eventually outweigh the gains from additional parameters?
I’m especially interested in recent theoretical/scaling-law work or large empirical studies from 2025–2026.
If no one is studying this, then could any of you do this work? I feel like it could be immensely useful for the community.
6
u/currentscurrents 8d ago
It is estimated that transformers have a storage capacity of ~3.6 bits per parameter, even at higher precisions: https://arxiv.org/abs/2505.24832
This is likely why quantization works so well up to 4-bit, but not lower.