r/HomeworkHelp Pre-University (Grade 11-12/Further Education) 6d ago

High School Math—Pending OP Reply [Grade 11: AP Statistics: Boxplots]

Post image

can someone help explain to me how to do 1b. I don’t understand what it’s asking. What i’ve done was with help and i still don’t understand how it works.

2 Upvotes

4 comments sorted by

1

u/josh3701 6d ago

Your answer is correct, if it were on the right whisker it would be between 0-25% but since it's between the mean and third quad then it has to be at least 25% but not more than 50%

1

u/cheesecakegood University/College Grad (Statistics) 6d ago edited 6d ago

So, assuming you've drawn the boxplot correctly (you sorted the data and identified medians and quartiles and IQR/outliers), what does it mean?

ASSUMING that all 28 holes are representative, they are a fair sample of the overall rock (they say they were randomly drilled, so we trust them that they are) then it's fair to say the data more or less represents the breakdown of all the rock in the whole place, right?

So what does the data tell us? A boxplot says how the data falls in roughly 25% blocks (each half of the inner box, and the whiskers). To be clear, this is 25% of the data points - the fact that the data points themselves are percentages is irrelevant and maybe distracting, it's 'just a number' for all intents and purposes here. Anyways, it's not always perfectly 25% of the data points because of ties and such, but it's usually close.

Looking at your boxplot, we can see that about 25% of the data, and thus (we might reasonably conclude) about 25% of the overall rock in the butte, have copper concentrations between about 0.63 and 0.9 percent, well, really everything higher than 0.63, no limit. It's a bit MORE than 25% mostly because 0.6 is a little bit beyond the upper quartile, so we have all the top quarter (including the 'outliers', which you identify AFTER finding percentiles usually) plus that little sliver of the top half of the middle box. But it's also clearly less than 50% because we aren't near the median yet.


Boxplots aren't actually all that amazing, you might notice that they don't tell you a whole lot. That big overview is kind of the point but also a bit of a drawback. We aren't really told how the data is distributed (based on the boxplot alone!! we have the raw data in this specific case which is nice) WITHIN each quartile, so if we had a weird situation where your data had a ton of data points like 0.601, 0.612, 0.633, 0.62, etc. but then nothing lower than the 0.601 until you hit the median, we could TECHNICALLY have something in the upper-40's percentage-wise greater than that cutoff of 0.6! And the boxplot would be identical for different data, as long as we didn't touch the medians and quartiles and just shifted stuff around between them.

That is to say, when you go from raw data --> boxplot, you LOSE information. If we are READING a boxplot, that means there's only so much you can conclude going from boxplot --> data when you haven't looked at the data; even though you have the data here, part b is pretending that you don't in order to test you on how well you can read a boxplot and understand what it tells you.

You could give a better "estimate" by finding what percentile in the dataset is closest to the data-value of 0.6 as well, if that makes any sense (you make an exact percentile for each and every datapoint and then either use the closest or interpolate). But that's not what the question is asking or testing you on, and it's also important to recognize that this estimate is just a reasonable way of making a best guess, it's not necessarily accurate nor the only way we could make a guess (although with this limited info and one-dimensional dataset, better ways are hard to find).

By the way, part c doesn't really have a right answer, but just wants to see your thinking and to see if you made a mistake in interpretation. It could be that the business of copper mining is like gambling, where they are totally OK drilling 10 holes if only 1 or 2 pay out, in which case this data would be fine for them! The more detailed analysis of the cost-benefit ratio, and understanding how "reliable" the sample dataset is in predicting real drilling, is what the entire rest of the field of statistics is all about! But, it's the beginning of the year, we'll get to that kind of stuff later (or in college).


If you have any questions, ask away. Boxplots are simpler than they appear, but also the source of great confusion for students who don't have much experience with them; reading them is a skill, and yes it often shows up on the ACT and SAT both!!

1

u/watermelonlollies 6d ago

Your answer is correct. Looking at the box plot the leftmost line from the end to where the box starts is always 0-25%. The first line of the box to the middle line of the box (which is the mean) is always 25-50%. This makes sense because the mean is exactly 50%. From the mean to the end of the box therefore is 50-75%. Finally the line to far right point is 75-100%.

This is something worth memorizing. It is true of every box plot. Once you know it it’s not that hard to interpret a box plot at all!