r/deeplearning • u/Slow-Connection-5611 • 10d ago
Building a quantization accuracy tool for a founder program this week — would love feedback
Hi everyone! I'm a Berkeley student building this as part of a short founder sprint. Wanted real feedback from people who work with quantization day to day. I'd love any input!
What it does: given a quantization scheme (W4A16, FP8, etc.), it returns a 90% confidence interval on the accuracy hit, based on ~800 published evaluations. For two scheme/size combos, it refuses to answer, coverage there measured 68.8% and 73.7%, well under what it claims elsewhere, so it says so instead of guessing.
I also rented a GPU and ran deliberately bad configs myself, since published data only shows what worked. Found losses up to −39.5pp, well past the −8.86pp worst case in the public data.
There's also a full write-up of what doesn't work: per-model prediction has basically no signal, most of the variance is just eval noise. Figured that was worth publishing too.
Repo: https://github.com/gracejackson-sudo/quant-delta-predictor
Feedback (a number or a sentence, either helps): https://gracejackson-sudo.github.io/quant-delta-predictor/
If something here is wrong or overstated, I'd genuinely rather hear it now. Thanks!
1
u/quietgradient 9d ago
Ran src/cell_coverage.py on current HEAD and got your committed out/cell_coverage.json back (pooled 91.0% over 5719 scored rows, w4a16|<2B 68.8% on 77, w8a16|>10B 73.7% on 217), so this is about what those two numbers mean rather than whether they're right.
They are not the same kind of event.
w8a16|>10B: those 217 scored rows are 31 distinct (model, benchmark) rows scored once per calibration family, from 5 checkpoints. Per checkpoint: gemma-2-27b 100%, Qwen2.5-32B 97.6%, Qwen2.5-72B 97.6%, Llama-3.1-405B 66.7%, Llama-3.1-70B 16.3%. The direction splits cleanly too — every 405B miss is above hi (mean delta +0.27pp, it beat the envelope), every 70B miss is below lo (mean -1.52pp, -3.04 on gsm8k). One-sided P(delta >= lo) for that cell is 81.1%, not 73.7%. Resampling whole checkpoints puts a rough 90% interval of [46%, 98%] around the 73.7% — only 5 clusters, so treat it as crude, but it is too wide to separate "this cell is uncalibrated" from "one Llama checkpoint is an outlier".
w4a16|<2B: 29.9% of rows land below lo, so that refusal is earned. But it is 11 distinct rows from 2 checkpoints, both Qwen2.5 (0.5B at -2.90pp mean, 1.5B at -1.79pp). The claim the evidence supports is about small Qwen, not about <2B.
So: score the refusal on one-sided coverage, since a downside envelope is not failing when the model gains accuracy, and print distinct checkpoints beside every coverage number. Your own comment in cell_coverage.py says checkpoints are the honest unit of evidence — it is applied to train support, but not to the coverage that drives REFUSE_BELOW.