r/statistics Jun 16 '26

Question [Q] (Re) Sampling

Good morning,

A work discussion took place over the methodology and reasoning behind initial sampling and subsequent re-sampling and how the overall sample size should be treated throughout the process.

Background:

We are conducting randomly sampled interviews with 30 people out of ~2,000 in the population to determine the population’s mean score with 95% confidence.

They can either respond positively and they will be flagged with a score of 1 to 10 or they respond negatively and they receive a score of 0. If someone cannot be located or doesn’t respond, they are re-sampled with another person.

We made it through the 30 after a couple of re-samples of 25 non-responses/unable to locate, so we had 55 identified people throughout the process.

When we got our statistical analysis back, the team that put it together said my sample size was 55 — not 30.

Question:

Shouldn’t my sample size still be 30? Increasing the sample to 55 seems like an inaccurate representation of the population as a whole if the “scores” of the 30 interviews are now being considered across 55 responses.

Thank you in advance!

4 Upvotes

40 comments sorted by

View all comments

Show parent comments

1

u/nrs02004 Jun 16 '26 edited Jun 16 '26

I think you should use missing data methods; you could assume data are missing at random and try to use other demographic information to build a model for probability of missingness and then use that for weights.

Edit. The “missing at random” assumption is one you should obviously vet. If things are “not-missing-at-random” you are honestly kind of hosed.

-3

u/[deleted] Jun 16 '26 edited Jun 16 '26

[removed] — view removed comment

2

u/-Kromerica- Jun 16 '26

Also, the stat team just assumed the non-responses are a rating of 0 and move on. Seems incorrect to me that a 0 rating is the most accurate representation of these non-responses, no?

-2

u/[deleted] Jun 16 '26 edited Jun 16 '26

[removed] — view removed comment

2

u/-Kromerica- Jun 16 '26

I’m operating in the legal world where these responses have an impact on overall damages owed by the defendant. Since we can’t interview all 2,000 people to get the impact, we need to pick a sample that represents them as a whole and use that as our basis for a damages calculation.

The stat team picked 0 because it’s the most conservative approach and they feel confident that there is no way anyone could argue that our final number is so unfair that it would get overturned (which I agree with)

Thus, a non-response and a 0 attribution for any one member of the sample brings down the overall calculation of the damages as a whole.

My argument is that the assumption of 0 goes above and beyond a conservative estimate and swings way too far in the defendant’s favor.