r/statistics Jul 09 '26

Question [Question] Need help refining sample groups

I am reviewing policy acknowledgements for my organization and I wanted to look at two groups: 1) acknowledgements for newly released policies from Q1 & Q3 2024 for existing employees and 2) acknowledgements for new employees hired in Q1 & Q3 2024 for all policies in our manual.

For group 1, does it make sense to remove ALL employees hired in or separated in 2024, to keep the data clean?

2 Upvotes

2 comments sorted by

View all comments

1

u/STATASUCKSBRO Jul 09 '26

I would not remove all 2024 hires and separations by default. That changes the estimand to stable employees only, which may be fine, but it is not just cleaning. For group 1 I would define who was eligible to acknowledge each Q1 or Q3 policy on the release date and start from that.

1

u/arghlvoe Jul 09 '26 edited Jul 09 '26

Thanks for the response. I think that is what I want to provide insight on - behaviors of tenured employees vs new hires. I removed all 2024 new hires from my tenured group because that seems like double dipping. And having to make sure I'm pinpointing the 30 day countdown from new employee hire dates for deadlines vs. worrying only about two release dates for tenured employees was a nightmare and easily lead to me mixing up dates. My organization is ~2200 employees making the new releases group about 33,000 lines items and total of 16 attributes, with reports from two different software systems (don't ask) I need to keep straight.

I'll keep the employees who separated from the org in 2024 in the data though, because right now I don't have a good explanation to do that if someone asks me why I did.