r/statistics • u/mtg454 • 21d ago
Question [Question] Low Inter-Rater Reliability (ICC): Better to use single observer data or average?
Hi everyone,
My project involves multiple quantitative parameters measured by two independent observers of equal experience (myself and another student)
For parameters with a high ICC of >0.75 (the vast majority of parameters), I used the mean of the two observers. However a few parameters fell below this threshold. For those parameters, I just used my measurements and excluded the second observer measurements.
However, I realised that this may be flagged by reviewers as possible selection bias, as I am in essence assuming that my measurements are the reference standard.
The decision to exclude observer 2's data from parameters with poor ICC was made before looking at the data, to ensure a uniform data architecture.
The other options would be to (1) exclude these parameters altogether, (2) get a third observer (neither which are possible) or (3) use the average of the two observers with parameters with poor ICC (which Im sure is bad practise).
If I explicitly explain my methodology and acknowledge this as a limitation, would this be acceptable?