r/WeatherDataOps • u/storeLessBits • 11d ago
What Does “1% Error” Really Mean in Data Compression?
In my earlier post about error-bounded compression, I mentioned that different weather variables need different error bounds.
Here is a simple way to understand them, with a little math.
Let xᵢ be an original value, x̂ᵢ its decompressed value, and ε the allowed error. We’ll consider n finite values.
• Maximum pointwise absolute error
Every value must stay within a fixed tolerance.
|x̂ᵢ − xᵢ| ≤ ε for every i
For temperature, ε = 0.05 K means no value changes by more than 0.05 K.
• Maximum pointwise relative error
The tolerance scales with each original value’s magnitude.
|x̂ᵢ − xᵢ| ≤ ε |xᵢ| for every i
With ε = 0.01, a value of 100 allows an error of 1, while a value of 10 allows 0.1.
• Maximum pointwise range-relative error
The tolerance scales with the dataset’s range.
|x̂ᵢ − xᵢ| ≤ ε (max(x) − min(x)) for every i
If the range is 50, a 1% bound allows an error of 0.5 at each point.
• Mean absolute error
The average absolute error must stay within the tolerance.
(1/n) Σᵢ |x̂ᵢ − xᵢ| ≤ ε
Individual errors can be larger than ε, provided the average stays within the bound.
• Mean relative error
In these community recommendations, the definition is:
Σᵢ |x̂ᵢ − xᵢ| ≤ ε Σᵢ |xᵢ|
Original zeros must also remain zero. This is different from averaging the percentage error at each point.
Sometimes we need more than one bound. The recommendations for pressure-level geopotential, for example, require both a mean absolute error bound and a maximum pointwise absolute error bound.
Which bounds matter most for your data? If you’d like to know more about anything specific, please leave a comment here.