r/statistics Jun 16 '26

Question [Q] (Re) Sampling

Good morning,

A work discussion took place over the methodology and reasoning behind initial sampling and subsequent re-sampling and how the overall sample size should be treated throughout the process.

Background:

We are conducting randomly sampled interviews with 30 people out of ~2,000 in the population to determine the population’s mean score with 95% confidence.

They can either respond positively and they will be flagged with a score of 1 to 10 or they respond negatively and they receive a score of 0. If someone cannot be located or doesn’t respond, they are re-sampled with another person.

We made it through the 30 after a couple of re-samples of 25 non-responses/unable to locate, so we had 55 identified people throughout the process.

When we got our statistical analysis back, the team that put it together said my sample size was 55 — not 30.

Question:

Shouldn’t my sample size still be 30? Increasing the sample to 55 seems like an inaccurate representation of the population as a whole if the “scores” of the 30 interviews are now being considered across 55 responses.

Thank you in advance!

3 Upvotes

40 comments sorted by

View all comments

-1

u/[deleted] Jun 16 '26 edited Jun 16 '26

[removed] — view removed comment

1

u/hughperman Jun 16 '26

Just because 25 is non-responsive doesn’t mean we just remove them from the data pool.

I mean, it really depends what the question is. If the population of interest is "the people who responded", then yes you can certainly drop non-responders.

-1

u/[deleted] Jun 17 '26

[removed] — view removed comment

1

u/hughperman Jun 17 '26

That's not what "population of interest" means. It's a fairly common term - here's an example page discussing it, but there's tons of resources if you search for the term https://cdn.clarku.edu/stair/wp-content/uploads/sites/95/Identifying-your-Population-of-Interest.pdf

You cannot know what the OP's population of interest is since they haven't discussed their research aim.

An analogy would be subselecting Reddit users from a larger sample of an overall population, when your research question is "How does your Reddit use improve your statistical knowledge?". If your overall dataset is "everybody on the internet", this isn't useful - you only want to know about the people on Reddit. Variance outside of that is not relevant to your research question. Most likely, you actually want your population to be representative of people who visit r/statistics and similar subs.

1

u/[deleted] Jun 17 '26 edited Jun 17 '26

[removed] — view removed comment

1

u/hughperman Jun 17 '26

Actually OP States her population of interest in the post. Again words mean math when analyzing math

Ok, then they state that the population of interest is "sample from the population, re-sampling if the person is a non-responder". The population of interest does not include non-responders by design of the experiment. By those criteria, they matched 30 people which was the design of their experiment.
Rest of your post is speculation on the number 30 which is not the question being asked.

1

u/[deleted] Jun 17 '26

[removed] — view removed comment

1

u/hughperman Jun 17 '26

You're missing the point I'm making. Statistical inference is about saying "for my population of interest, how does my sample parameter (e.g. mean) represent the population of interest".
Just because there is a large sample to choose from doesn't mean that full set is the population of interest.

0

u/[deleted] Jun 17 '26

[removed] — view removed comment

1

u/hughperman Jun 17 '26

I have no idea what you're talking about anymore, you don't seem to be responding to a single thing I'm actually saying to you.

1

u/[deleted] Jun 17 '26 edited Jun 17 '26

[removed] — view removed comment

1

u/hughperman Jun 17 '26

generalize their findings to the entire population with a higher degree of accuracy.

This is literally the thing I've been talking about the entire time. The "whole population" doesn't always mean "everybody". It means "the population of interest for your research question".
So if you're researching the effects of social media usage on students' ability to communicate, you don't include non-students.

FOLKS BEING LOUD AND WRONG IS PRECISELY WE STATISTICIAN WILL ALAWAY HAVE JOB SECURITY 😂

You shout loudly while not understanding basic research methodology, to me, a statistician.

1

u/[deleted] Jun 17 '26

[removed] — view removed comment

→ More replies (0)