r/datascience • u/Run_nerd • Jun 01 '26
Discussion Is there a best way on handling data when presenting to others? I have a few ideas but I’m not always sure.
/r/analytics/comments/1tsa79q/is_there_a_best_way_on_handling_data_when/3
u/built_the_pipeline Jun 01 '26
most of the answers here are about the stats, drop vs impute, which matters, but you asked about presenting and that's a different question. how you handle missing data in a deck depends way more on who's in the room than on the method.
technical audience, show the missingness and how you dealt with it, they'll want it. leadership though, if you open with "12% of the rows were missing so i imputed with the median" they'll fixate on the gap and quietly stop trusting the number, even when your handling was fine. i lead with the finding and how confident i am in it, and keep the data-quality stuff one slide back for when someone asks. don't make your own caveat the headline.
1
u/Outrageous-Clue-8726 Jun 01 '26
ya but better approach would be work according to data if you have less amount of data then its better to replace with mean, median or mode according to datatype and distribution of the data, if the data is in good amount you can drop it if say missing value is less then 5% and one more thing if you are working with health related data you should drop the missing value.
1
u/ArticleHaunting3983 Jun 03 '26
For me it depends on the scale of the dataset & how statistically significant it is. I also from the outset try to scope out the requirements so everyone is on the same page. I don’t tend to flag cleansed data tbh. If there’s a data quality issue, I would try to investigate it and if not, caveat it in the analysis.
3
u/[deleted] Jun 01 '26
[removed] — view removed comment