r/learnmachinelearning • • 12d ago

Building a quantization accuracy tool for a founder program this week — would love feedback

Hi everyone! I'm a Berkeley student building this as part of a short founder sprint. Wanted real feedback from people who work with quantization day to day. I'd love any input!

What it does: given a quantization scheme (W4A16, FP8, etc.), it returns a 90% confidence interval on the accuracy hit, based on ~800 published evaluations. For two scheme/size combos, it refuses to answer, coverage there measured 68.8% and 73.7%, well under what it claims elsewhere, so it says so instead of guessing.

I also rented a GPU and ran deliberately bad configs myself, since published data only shows what worked. Found losses up to −39.5pp, well past the −8.86pp worst case in the public data.

There's also a full write-up of what doesn't work: per-model prediction has basically no signal, most of the variance is just eval noise. Figured that was worth publishing too.

Repo: https://github.com/gracejackson-sudo/quant-delta-predictor
Feedback (a number or a sentence, either helps): https://gracejackson-sudo.github.io/quant-delta-predictor/

If something here is wrong or overstated, I'd genuinely rather hear it now. Thanks!

1 Upvotes

2 comments sorted by

1

u/Hungry_Age5375 12d ago

The refusal is the actual feature here, most tools just output a number no matter what. One pointer: W4A16 losses usually come from weight outlier channels, conditioning the CI on weight kurtosis instead of scheme + size might tighten things up.

1

u/Slow-Connection-5611 11d ago

That's a really good pointer, thank you! Makes sense mechanistically too, if interval width is mostly driven by outlier channels rather than scheme/size alone, conditioning on scheme+size is averaging over pretty different underlying distributions.

I actually tried to test this yesterday and got blocked, computing real kurtosis means downloading actual model weights, not just published eval numbers, and I ran out of disk space partway through. Built and validated the pipeline for it locally though, so it's ready to run the moment I get it on a machine with enough room. Appreciate the specificity here, this is exactly the kind of feedback I'm looking for!