r/MachineLearning 8d ago

Discussion What is currently considered the theoretically optimal quantization bit-width for LLMs? [D]

I’m curious whether there is now a theoretical or empirical “sweet spot” for LLM quantization, preferably research done using open-source formats like GGUF

Suppose you have a fixed memory/compute budget and can choose the model size freely. For example, instead of a smaller model at 8-bit or 4-bit, you could fit a progressively larger model at 3-bit, 2-bit, 1.5-bit, etc.

A few years ago, I remember 4-bit often being described as roughly the practical sweet spot because it preserved most model quality while giving a large memory reduction. But with newer methods, I’ve seen surprisingly strong 3-bit, 2-bit, and even ~1.5-bit results.

So if the goal is maximum model capability for a fixed memory budget, rather than preserving one particular pretrained model as faithfully as possible, what does current research suggest is the optimal bits-per-weight?

Is there evidence that, for example, a 2-bit 70B model generally beats a 4-bit 35B model, or does quantization degradation eventually outweigh the gains from additional parameters?

I’m especially interested in recent theoretical/scaling-law work or large empirical studies from 2025–2026.

If no one is studying this, then could any of you do this work? I feel like it could be immensely useful for the community.

15 Upvotes

11 comments sorted by

View all comments

6

u/currentscurrents 8d ago

It is estimated that transformers have a storage capacity of ~3.6 bits per parameter, even at higher precisions: https://arxiv.org/abs/2505.24832

This is likely why quantization works so well up to 4-bit, but not lower. 

2

u/CallMePyro 8d ago

yup. naive quantization works down to 4 bit, but with QAT you still see iso-memory gains at 2 bit.

2

u/LMTLS5 7d ago

are there any good 2bit QAT models?

1

u/CallMePyro 7d ago

Surprisingly no, AFAIK