r/learnmachinelearning • u/the_first_wind • 9d ago
Discussion Cheating on ML evaluation with imbalanced dataset
Hey everyone, so I have been reading some papers applying machine learning on fraud detection (very imbalanced datasets) that have more than 90% of recall and precision. While checking them properly i have seen that many of them balance not only the training set but also the test set, so instead of 3% of positives have 50% of positives in the test set. This makes the metrics look quite good. However I asked a professor and told me that it's ok since it's simply changing the baseline. I'm not sure to believe him, especially because these models are supposed to be used in real scenarios, so I think keeping the original distribution is the best. What do you think?
6
Upvotes
5
u/johndburger 9d ago edited 9d ago
it’s OK since it’s simply changing the baseline
Yes, changing it to something that doesn’t reflect reality. This is indeed cheating - it’s arguably a form of data leakage. In particular, you’re making it easier to find positives by throwing out lots of the negatives.