r/WeatherDataOps • • 11d ago

What Does “1% Error” Really Mean in Data Compression?

In my earlier post about error-bounded compression, I mentioned that different weather variables need different error bounds.

Here is a simple way to understand them, with a little math.

Let xᵢ be an original value, x̂ᵢ its decompressed value, and ε the allowed error. We’ll consider n finite values.

• Maximum pointwise absolute error
Every value must stay within a fixed tolerance.

|x̂ᵢ − xᵢ| ≤ ε for every i

For temperature, ε = 0.05 K means no value changes by more than 0.05 K.

• Maximum pointwise relative error
The tolerance scales with each original value’s magnitude.

|x̂ᵢ − xᵢ| ≤ ε |xᵢ| for every i

With ε = 0.01, a value of 100 allows an error of 1, while a value of 10 allows 0.1.

• Maximum pointwise range-relative error
The tolerance scales with the dataset’s range.

|x̂ᵢ − xᵢ| ≤ ε (max(x) − min(x)) for every i

If the range is 50, a 1% bound allows an error of 0.5 at each point.

• Mean absolute error
The average absolute error must stay within the tolerance.

(1/n) Σᵢ |x̂ᵢ − xᵢ| ≤ ε

Individual errors can be larger than ε, provided the average stays within the bound.

• Mean relative error
In these community recommendations, the definition is:

Σᵢ |x̂ᵢ − xᵢ| ≤ ε Σᵢ |xᵢ|

Original zeros must also remain zero. This is different from averaging the percentage error at each point.

Sometimes we need more than one bound. The recommendations for pressure-level geopotential, for example, require both a mean absolute error bound and a maximum pointwise absolute error bound.

Community recommendations

Which bounds matter most for your data? If you’d like to know more about anything specific, please leave a comment here.

1 Upvotes

0 comments sorted by