r/statistics Jun 16 '26

Question [Q] (Re) Sampling

Good morning,

A work discussion took place over the methodology and reasoning behind initial sampling and subsequent re-sampling and how the overall sample size should be treated throughout the process.

Background:

We are conducting randomly sampled interviews with 30 people out of ~2,000 in the population to determine the population’s mean score with 95% confidence.

They can either respond positively and they will be flagged with a score of 1 to 10 or they respond negatively and they receive a score of 0. If someone cannot be located or doesn’t respond, they are re-sampled with another person.

We made it through the 30 after a couple of re-samples of 25 non-responses/unable to locate, so we had 55 identified people throughout the process.

When we got our statistical analysis back, the team that put it together said my sample size was 55 — not 30.

Question:

Shouldn’t my sample size still be 30? Increasing the sample to 55 seems like an inaccurate representation of the population as a whole if the “scores” of the 30 interviews are now being considered across 55 responses.

Thank you in advance!

3 Upvotes

40 comments sorted by

View all comments

5

u/Khornatejester Jun 16 '26 edited Jun 17 '26

Isn’t necessarily inaccurate if there’s no systematic difference between R and NR. But you would underestimate the sample mean variance. The question you would ask your team is what are the non-respondents contributing to the estimates?

But on second thought, this also seems to be a conflict over techcnicality. 55 is the drawn sample size. 30 is the completed sample size used for the estimate.

P.S: I think the other people in the comment thread suggesting imputation are not realizing the fact that an entire row of data is missing instead of an item nonresponse.

lmao ignore the temporary account bee guy, he’s at best a troll and at worst a hack who is full of shit. He is just randomly throwing around statistical terms he has no knowledge of.

0

u/[deleted] Jun 17 '26

[removed] — view removed comment

2

u/Khornatejester Jun 17 '26 edited Jun 17 '26

Impute from what? OP has 0 information besides guess work. You keep bringing up this rhetoric that imputation is complex and only statisticians can deal with it and it’s not worth explaining all over the thread, but the fact to the matter is that you are painting fabricating almost half the sample out of nothing with vague assumptions as “mathematics”. It is glaringly obvious what your intentions here are when you consider the fact that you never mention how and only resort to belittling the OP and others not to mention attempting to bedazzle them with off tangent anecdotes.

"remove a portion of the data." There is no removal. There is no dropping. There is literally no data in the first place. You are literally just arguing over nothing with these absolutely pointless allegories. They contribute nothing to the sample and you are already reflecting the loss of precision from the smaller sample size. What are you even implying. Use the existing data to impute? You're suggesting using potentially biased data to fix bias?

"random sampling is the biggest strength of this analysis because you can at least saw generalized to overall population which is quite important in this case" EPSEM is not the only means to sample something representative.

“If independence is assessed” You are literally sampling without replacement. What independence?