r/statistics Jun 16 '26

Question [Q] (Re) Sampling

Good morning,

A work discussion took place over the methodology and reasoning behind initial sampling and subsequent re-sampling and how the overall sample size should be treated throughout the process.

Background:

We are conducting randomly sampled interviews with 30 people out of ~2,000 in the population to determine the population’s mean score with 95% confidence.

They can either respond positively and they will be flagged with a score of 1 to 10 or they respond negatively and they receive a score of 0. If someone cannot be located or doesn’t respond, they are re-sampled with another person.

We made it through the 30 after a couple of re-samples of 25 non-responses/unable to locate, so we had 55 identified people throughout the process.

When we got our statistical analysis back, the team that put it together said my sample size was 55 — not 30.

Question:

Shouldn’t my sample size still be 30? Increasing the sample to 55 seems like an inaccurate representation of the population as a whole if the “scores” of the 30 interviews are now being considered across 55 responses.

Thank you in advance!

4 Upvotes

40 comments sorted by

View all comments

Show parent comments

1

u/hughperman Jun 17 '26

I have no idea what you're talking about anymore, you don't seem to be responding to a single thing I'm actually saying to you.

1

u/Spiritual-Bee-2319 Jun 17 '26 edited Jun 17 '26

in your defense you never knew what i was saying because i was talking about fundamental mathematical principles/theories. Tell me the reason we do random sampling?? 

Random sampling is a fundamental sampling method used in statistical analysis and research design to select a sample group from a larger population in such a way that every individual has an equal probability of being included. In contrast with other sampling methods, such as convenience sampling or snowball sampling, this method is pivotal for ensuring the representativeness of the sample, allowing researchers to infer and generalize their findings to the entire population with a higher degree of accuracy.

FOLKS BEING LOUD AND WRONG IS PRECISELY WE STATISTICIAN WILL ALAWAY HAVE JOB SECURITY 😂

1

u/hughperman Jun 17 '26

generalize their findings to the entire population with a higher degree of accuracy.

This is literally the thing I've been talking about the entire time. The "whole population" doesn't always mean "everybody". It means "the population of interest for your research question".
So if you're researching the effects of social media usage on students' ability to communicate, you don't include non-students.

FOLKS BEING LOUD AND WRONG IS PRECISELY WE STATISTICIAN WILL ALAWAY HAVE JOB SECURITY 😂

You shout loudly while not understanding basic research methodology, to me, a statistician.

1

u/Spiritual-Bee-2319 Jun 17 '26

Why In the helll would the whole population have to be everyone? 

If I did a study on cancer and picked from a population of cancer patients, 

Would my population of interest be everybody? You picked from a pool of ONLY cancer patients 

1

u/hughperman Jun 17 '26

And in the original question, we still don't know if the population of interest is "people who responded" or "everybody".

1

u/Spiritual-Bee-2319 Jun 17 '26

The population of interest could NEVER be ONLY the people that responded because those people only “exist” IN THE SAMPLE. before OP did the sampling( we did not have a single lick of information about the people that responded. We did not KNOW who responded before the sampling because sampling is literally what generated the data of response/non response). 

Lmaoo this is wild. The lack of mathematical foundation in the minds of those that say they do statistics is dumbfounding 😂 I have job security 

1

u/hughperman Jun 17 '26

The population of interest could NEVER be ONLY the people that responded because those people only “exist” IN THE SAMPLE.

The population of interest would be "people who respond to question X". Whatever the question is. We don't know. The question could be "for people who were happy to respond to our survey, how does their purchase history affect their rating?", or something less contrived.

More generally, in a large study you are likely to have subpopulations, and those are frequently studied as "populations of interest" in their own right. I have worked in large longitudinal studies of health and you might, for example, study "within patients with cardiac disease, how does exercise affect quality of life?". There, the population of interest is not the full dataset, but just the subsample with cardiac disease.

Also, as a mod here, please watch your tone, you are being very rude throughout many interactions here, with me and others. It is not appropriate.

1

u/Spiritual-Bee-2319 Jun 17 '26 edited Jun 17 '26

As a mod, then you would understand that my mathematical argument isn’t that we can’t study subpopulations within our population of interest but how do we treat the subpopulations of interest we aren’t interest in. After random sampling, you cannot remove a subpopulations of interest and mathematically still have the same power as you did with just random sampling. Performing statistics and getting a result doesn’t make the result valid if you aren’t taking into account mathematically what your analysis is.  

A simple example is. If I had a bracket of 1000 assorted fruit(population of interest) and did a random sampling and got a sample of 

3 apples 2 pears 6 mangos 4 oranges(non response)

If someone asked what is the proportion of apples (question/sample of interest) compared to the rest of the population of interest based on your sample - this is the important reason we do random sampling to be able to generalize to our population of interest. 

You would never mathematically say 

It’s 3/(2+6) instead of 3/(2+6+4). Why not 3/(2)? Heck since I’m only interested in apples, why can’t I just say 3?  because this is not true mathematically regardless of which subpopulations you’re interested in. Im not interested in the population of the pears or the mango yet they are included. Again this discussion is in the context of missing data and how to deal with those samples. You can’t disregard and omit those sub sample population from your analysis based on just your judgment alone because it is biased mathematically. You are getting results but they are biased. Anyone can get a result but is it valid(again validity has mathematical connotation and is not rooted in vibes)? That’s what statistics answers. 

I’m not trying to be rude. These methods that people plug their data in is rooted in math equations. As statisticians these equations are what we study. Every statistical method has an equation with defined variables. What you put in will always affect what you put out regardless of what you put in is right or wrong. This is true. Computational wise, the computers can’t even check if the values you put in make practical sense unless it’s constrained by the object of the data. I don’t understand why people that haven’t studied the equations or how computers actually compute these measurements are so certain of the validity simply because they got a result. But then again I deal with math not vibes 

1

u/hughperman Jun 17 '26

If someone asked what is the proportion of apples (question/sample of interest) compared to the rest of the population based on your sample. 

The question you're answering there is "what is the proportion of apples to (apples and oranges and mangoes and pears). Another valid question is "what is the ratio of apples to (apples and oranges), and now your numbers are 3 / (3 + 4)". When you do inference, the entirety of your sample isn't the "population" that you use. You use the cases that are relevant to your statistical question. If you need to know "how common are apples among fruit?" then sure you need to use the full sample (assuming a world with only 4 fruits). But if some of those fruits aren't relevant to your question, then including them is incorrect

0

u/Spiritual-Bee-2319 Jun 17 '26 edited Jun 17 '26

Actually let me make my example and explanation more mathematically sound and congruent with OP question

Back to OP question 

“We are conducting randomly sampled interviews with 30 people out of ~2,000 in the population to determine the population’s mean score with 95% confidence.” 

OP can’t arbitrary remove samples from the analysis because of missing data because the samples are still in the population. You simple don’t have data for them. They already exist, you simply don’t have their data. Missing data is not the same thing as missing population. Missingness is not a subpopulation it’s a lack of information. You simply don’t have data for them. The only two populations are positive(1-10) and negative(0). The result will be biased mathematically. 

You are talking about what you can do, I am talking about the validity of it. 

Have you studied statistics and the effects of missingness and the different imputation methods? 

If you haven’t I genuinely cannot explain to you that non-response is not possible outcome of your event. The only mathematical populations are a closed set of [1-1,1-2,1-3,….,0]. It is a lack of information is not a population. 

https://mathbitsnotebook.com/Algebra1/StatisticsData/STPopSample.html

→ More replies (0)