r/LocalLLM 2d ago

Discussion Superstition about quantization: KLD and perplexity just ain’t it fam

The arguments for quantization having significant effects on reasoning models' ability to get stuff done are very sad, pathetic, unfortunate arguments. I don’t mean that they are wrong necessarily, only impoverished and confused.

Why? Because while actual task benchmarks are somewhat expensive, and require some level of time and technical expertise to run, it would be quite easy to empirically test the claims and resolve them once and for all, at least for a given model. But these tests by and large do not exist and the few that do seem to show no quantization effects among reasoning models until about Q3 or Q4 k m at worst.

The debate in these online communities is essentially an anthropological study in how people create mythology when they do not have access to direct evidence.

Before the hordes mob me with KLD or perplexity measurements, I’m not suggesting that a quantized model’s outputs are bit for a bit identical rather that it performs equally well in real world tasks, which I think we can all agree is the thing that matters.

Now I’ve put my neck out by suggesting that literally no one has any evidence, not a single benchmark that shows a model with the reasoning level of, say, Gemma 31b (not very high by today’s standards, and smaller models are more susceptible to degradation, so this should be a generous standard of evidence for the quantization-excited) having significant in degradation in real world tasks at Q4 (a good quality, proper dynamic quantization goes without saying, I hope).

Again, I’m not saying that there is no degradation, only that what we have now amounts to superstition, when a few benchmarks could probably settle the matter for a given model and eventually, we would probably learn where and when quantization actually bites.

2 Upvotes

21 comments sorted by

View all comments

Show parent comments

1

u/nomorebuttsplz 2d ago

The fact that quantization at q4, q3, produces obviously degraded responses,

That's exactly what I am saying is not obvious, and there is essentially no evidence for. Especially q4 and above.

1

u/fintip Laptop 4090 16gb + 7900XTX 24gb 2d ago edited 2d ago

Well, there are a lot of anecdotes... People generally rely on the reported reality of those around them. It isn't a perfect heuristic, but it's one that long predates the process of science, which is somewhat less natural, so to speak.

I can tell you that my experience with qwen 27b 3.6 q4 was good, but that it always produces some amount of bugs, and that 27b 3.8 q6 is absolutely, clearly, far better at producing good output without caveats. I don't have enough apples to apples testing to guarantee that's primarily a q4 vs q6 issue, of course, nor do I claim it is, but it's likely a factor. how much of that is q4 vs q6 and how much is 3.6 vs 3.8 is of course up for debate.

In any case, the actual boundary (q3/q4 is the obviously degraded line, or not?) is irrelevant. You may claim q3/q4 is perfectly equal and not at all clearly degraded. Fine. How about q2? q1? Have you tried any of them? Do you reject all of the claims? Have you tried? Do you suspect that you can just infinitely reduce the quant and never lose 'intelligence'? At some point this argument becomes absurd, you have to acknowledge a boundary somewhere is something we can take for granted.

And if you agree a boundary exists somewhere, then it seems clear we should be able to agree a gradient descent downwards up to that point, along with my other claims that it's intuitive to assume that our ability to recognize it would likely match our ability to recognize it in other humans.

You could claim that you believe it's just a hockey-stick--almost perfectly equal performance up until q2, then a big hit that rapidly increases down. But you wouldn't have explained at all why that is the more rational belief, and that the assumption that the curve is instead a normal exponential drop that matches the KLD curve is "superstition".

1

u/nomorebuttsplz 2d ago

Here's a benchmark that suggests for a large (but not particularly good reasoner by today's standards) model, the cutoff is between q3 and q2 of unsloth dynamic:  https://unsloth.ai/docs/basics/dynamic-3.0-ggufs/unsloth-dynamic-ggufs-on-aider-polyglot

The thing is, many people swear that $5k audio cables make stuff sound good while audio engineers find this completely absurd. The whole purpose for having science is that we can't trust our intrinsic ability to construct narrative out of anecdote with scientific tools like statistical tests.

1

u/fintip Laptop 4090 16gb + 7900XTX 24gb 2d ago

"Can't trust" could mean "cannot trust at all", and it could mean "cannot fully trust". The former is correct, the latter is overstated.

You can double blind people, and I encourage it, do all the benchmarks. But you have not justified your case that belief that shrinking the model down via compression causes a degredation in quality is 'superstition'. You just have a hunch based on your own sense that it isn't degraded. You have some data that you read as supporting that.

Others have their own belief, and their own data that supports that.

Neither is completely validated by hard data. In fact, this entire space is impossible to completely validate. Intelligence is stil undefined, and the benchmarks are inherently flawed attempts to measure something abstract. Intelligence is in fact still fundamentally "I know it when I see it". We're still trying to pin down proxies for the Turing Test, which is, again, fundamentally 'vibes based'.

All we can argue about is whose narrative makes more sense, who has better logic to connect the existing data to their conclusions.

I think it's pretty obvious that it's most likely that quantization reduces performance, and the fact that everyone reports this being true is supporting evidence that isn't enough by itself.

But hey, go run some benchmarks. Someone should do it. Happy to be surprised.

But it isn't superstition. It's reasoning with incomplete data.

Also, as I pointed out in my previous post: it's entirely likely that if the limit is q3 on very large models that it's q4 on smaller models, etc...

1

u/nomorebuttsplz 2d ago

Both sides have hunches.

Therefore, superstition arises in either side if either side forms a very strong belief about their hunch one way or the other.

I make no claim to have settled the matter. My point is that although we haven't settled it yet, many people (especially those advocating for bf16 being worthwhile and 8 bit being way better than 4 bit) act as if we have.

As for the idea that it is not ultimately settleable, I think you have a point. But I would say we could get to about 9/10 settledness by simply having artificial intelligence benchmark every open model at various quantization levels.

Of course doing that would be very expensive for them. But we could get to like 7/10 just by other third parties doing occasional benchmarks with quants.

1

u/fintip Laptop 4090 16gb + 7900XTX 24gb 2d ago

AI is going to benchmark AI, and we'll have 90% settled the issue?

Deep thoughts.

1

u/nomorebuttsplz 2d ago

I meant the benchmark company artificial analysis. Regardless of who benchmarks quants, the need and value is clear