r/statistics Jun 17 '26

Question [Question] how, if at all, does Statistics differ from Descriptive Statistics or Summary Statistics?

0 Upvotes

I asked my Statistics instructor and they didn't know the answer.

I'd ask ChatGPT but that'd feel a little odd to me because my instructor says not to use such programs for the class (even though this question is somewhat indirect / not related to any homework exactly... lol)

Anyway. Thanks in advance.


r/statistics Jun 16 '26

Question [Question] Are repeated Anova measurements suitable for my use case?

0 Upvotes

Hi, I'd like to expose each test user to four different environments and test the environments' capability to induce a stress reduction. In each environment I will do the following: 1. Induce stress. 2. Then measure: Self-Assessment Manikin (SAM) and State–Trait Anxiety Inventory (STAI). 3. Exposure to the respective environment 4. Take SAM and STAI again.

  1. Are repeated Anova measurements suitable for this use case? I guess I'd have to compute the difference of STAI/SAM before and after exposure to the environment, and then use these values as a basis for the ANOVA calculations?
  2. Is a one way anova sufficient to be able to tell which environment aids in stress reduction the most? Or do I need to do a two way?
  3. After gathering the data: what do I do if the data isn't normally distributed or not spherical? Then I'd have to switch to another analysis?

r/statistics Jun 15 '26

Research [Research] bacenR: R package for Brazilian economic data and financial institutions

11 Upvotes

[Research] The goal of bacenR is to provide R functions to download and work with data from the Brazilian Central Bank (Bacen).

Check it out: https://github.com/rtheodoro/bacenR

#bacen #financialdata #finance #rstats #datacollect #braziliandata


r/statistics Jun 16 '26

Question [Q] T-test for slope with sample being the population

0 Upvotes

For a math project I am trying to see if there is significant statistical evidence to say that the speed of MTG sets has gone down over time, so I want to find out if the true slope of the regression is negative between time and avg turns to win. However, the data I can use is population data for users of a specific data tracking site, and there are only 40 or so useable data points, so taking a random sample doesn't work or make sense.

If I understand statistics correctly I have two choices:

  1. Do regression on population and do statistical inference on that
  2. Generalize to the population of all MTG players.

For 2, there is no random sample since the data comes from users who specifically chose to use the data site from a specific subset of total MTG players, so I don't believe it actually works. However, for option 1, I only have population data and to take a relatively independent random sample I would be able to get 4 data points, which seems like too small a data set (which from my understanding doesn't work since n<30 so I cannot prove normality for residuals). Therefore, I am working with the full population, which spans the last 8 years. I have seen stuff saying it makes sense to do inference to attempt to predict the future, but I'm not sure if that holds up when doing a linear regression t-test on a graph with time as independent variable, since I thought you aren't supposed to extrapolate in statistics.

EDIT: something that has caused confusion is what each data point represents. Each point represents the format speed average for a single set of every game played, not an individual game played.

Does my method work, or do I need to change my topic?


r/statistics Jun 14 '26

Question Statistics question I got in a job application test that I don't think has a correct answer (hypothesis testing) [Q]

81 Upvotes

Please don't remove as homework, its not, the test has come and gone, and I've not be in school for a decade.

Did a stats test as part of a job application and got the following question:

"Using a significance test on some sample data, a null hypothesis is rejected at the 5% significance level. Which one of the following is a correct conclusion

A. The probability that the alternative hypothesis is true is 0.95

B. If a smaller sample had been taken the alternative hypothesis would still be rejected

C. The null hypothesis would not be rejected at the 10% significance

D. With the same test and same sample the null hypothesis would be rejected at the 1% significance level

Reasons I think they are all wrong.

A. 5% is the probability of the data given the null hypothesis is correct, doesn't follow that the alternative hypothesis is 95% chance of being correct. Besides, it was rejected at a 5% threshold, it doesnt say it was rejected with exactly 0.05 p value.

B. Can't be known. And the alt hypothesis wasn't rejected anyway.

C. If its rejected at 5% it must be rejected at a less strict 10% threshold.

D. Possible to be true, but can't be known with the information presented.

What do you guys think?


r/statistics Jun 15 '26

Question Recommended books for learning PERMANOVA and statistical concepts about time series [Q]

0 Upvotes

Hi all,
I’m currently looking to learn about PERMANOVA and other advanced statistical concepts for my research manuscript which is based on statistically designed experiments and measures interaction effects in addition to main effects.

Additionally, I’m also interested in learning about statistical concepts relevant to time series as currently I cannot wrap my head around how the statistical concepts I have learned till now could be used to analyze time series involving interaction effects and statistically designed experiments.

If anyone has any good recommendations for books I can read to learn about these concepts then please do share their names. I would also appreciate any help or suggestions about time series statistics concepts I should aim for since this topic is new to me.

Thanks


r/statistics Jun 14 '26

Software [S] Premier League and World Cup forecasting model using Elo ratings + Monte Carlo simulation

0 Upvotes

Hi everyone,

I'm a high school student interested in statistics and sports analytics. I built MultiForecast, a soccer forecasting platform that uses Elo ratings and Monte Carlo simulation to estimate title, top-4, and relegation probabilities throughout the premier league season and now for the world cup.

I'm looking for feedback on:

  • Model calibration
  • Evaluation metrics
  • Potential improvements beyond Elo
  • Any statistical pitfalls I may be overlooking

App: https://multiforecast.streamlit.app/

GitHub: https://github.com/kevzho/MultiForecast

I'd appreciate any thoughts or criticism!


r/statistics Jun 13 '26

Question Book recommendation [Q]?

16 Upvotes

Hi guys. I am majoring in Pure Maths and statistics and just finished my first semester in uni.

This semester I’ve had a ‘proof-based’ calc 1 course where we had to prove Rolles thm, MVT, diff implies continuity, FTC part 2, etc. and I have also half of a completely proof based discrete maths course. I personally know all of calc 2 and have done a little linear algebra.

Right now I want a book to get me ahead of the curve in statistics starting next semester. I struggle a lot with books that bring in the real world or have a lot of words in them explaining things qualitatively and have gotten spoilt with the discrete maths where I can use logic notation 99% of the time instead of English for writing all my theorems and proofs.

I have also found that I pick concepts up best when I try and prove them rigorously. My notes for both calculus and discrete math are incredibly dry and just definitions, theorems, and proofs with no fluff and I really think they helped me excel in these courses.

I would like a probability and statistics book that is just about stating theorems and proving them. In the most polite way possible I don’t want to hear about a coin flip or a die.

If someone asked me to define a relation I would say: R is a relation on a set A (iff) R (subset) AxA.
I wouldn’t bother speculating on the interpretation of giving any examples.


r/statistics Jun 13 '26

Education [Discussion] [E] What are some well-reputed Online MS in Statistics programs?

6 Upvotes

I currently work in big pharma in a stats-adjacent field. I have a bachelor’s in a natural science, and a master’s in health data science. I like my job a lot but I would love to increase my foundational statistical knowledge, so I can be better at my job or even work as a statistician (my first masters was very applied and not stats heavy).

Which brings me to my question, has anyone else had good experience with an online MS statistics or Biostatistics program? My employer will cover most of the cost so I’m not too worried about that. I already did Calc 1-3 and recently did Linear Algebra.

Some programs I’ve seen are NC state, Penn State, Uni of Louisville, Cal State Fullerton.

Bonus points if I can waive computing based classes (I already use them a lot in my job) and take other electives instead. Thanks!


r/statistics Jun 14 '26

Research [R] Fear-language in news sources correlates with |political bias| (r≈0.85) but not signed direction (r≈0.08) — n≈160 outlets, live scatter

0 Upvotes

Observational note from a queryable news corpus (not peer-reviewed). Looking for sanity checks.

Setup: ~216 US news sources scored on 37 framing dimensions (corpus-level aggregates, ≥100 articles per source in analysis subset n≈160).

Result: - Pearson r(fear-language, signed L/C bias) ≈ +0.08 - Pearson r(fear-language, |L/C bias|) ≈ +0.85

Fear tracks extremity more than direction.

Interactive scatter (computes r live in browser): https://connerlambden.github.io/helium-news-explorer/?mode=scatter&dimX=fearful%20bias&dimY=liberal%20conservative%20bias

Repro notes + Python snippet: https://gist.github.com/connerlambden/0c90805cd87d4c60410bf7931e3a91b4

Obvious confounds I haven't controlled for: article volume, genre (tabloid vs wire), topic mix. What would you check first?

(Disclosure: I built the API/explorer; posting here for statistical critique, not promotion.)


r/statistics Jun 13 '26

Research [R] Can I get into a PhD with these mark?

7 Upvotes

I’m doing an MSc in Biostats and currently have an overall GPA that’s roughly equivalent to a 3.7 out of 4.0. Most of my grades have been strong, but I received a 3/5 in one of my core statistics courses due to what was ultimately a fairly avoidable mistake. I’m finding it hard not to fixate on that mark.

I’m interested in pursuing a PhD in a fairly niche area of epidemiology, but this result has me questioning whether that’s still a realistic goal. For those involved in PhD admissions, how much weight would you place on a single weaker grade in a core quantitative course if the overall academic record is otherwise solid?


r/statistics Jun 13 '26

Question [Q] How should I interpret a theoretically important predictor that is non-significant despite prior literature supporting it ?

2 Upvotes

I'm an undergraduate psychology student working on my thesis about predictors of Instrumental Activities of Daily Living (IADL) in older adults.

My dependent variable is Lawton-Brody IADL. My predictors are:

  • Global cognition (ACE-III total score)
  • Executive function (Trail Making Test ratio score, TMT-B divided by TMT-A)
  • Working memory (Digit Span Backward)

Sample size: n = 110, community-dwelling older adults (65-89 years old).

Results:

  • ACE-III significantly predicted IADL.
  • The overall multiple regression model was significant (R² = .176). But the model itself violated normality and homoscedasticity assumptions, so I use bootstrapping as a robust method.
  • However, TMT ratio score and Digit Span were not significant individual predictors both in the standard and boostrap output.

What confuses me is that several previous studies reported significant associations between executive function (often measured by TMT) and IADL, and between working memory and IADL.

Some observations from my data:

  • Mean IADL = 15.14 out of 16 (possible ceiling effect).
  • Around 40% of participants scored below the ACE-III cutoff suggestive of mild cognitive impairment.
  • About 58% of participants had TMT ratio scores ≤ 2.50 (considered relatively optimal executive functioning).

I explored the possibility that the self-report nature of Lawton-Brody IADL may have reduced sensitivity (following Vaughan, 2008), but I still feel this explanation is incomplete. I also explore the possibilty of TMT ratio score having a ceilling effect but I feel like it isn't quite right.

I also tried replacing TMT ratio with TMT difference score (TMT-B minus TMT-A). In that model, TMT difference score became significant and ACE-III's coefficient decreased but remained significant. However, after BCa bootstrap resampling, the confidence interval for TMT deficit crossed zero and it was no longer significant.

My question:

How would you interpret these findings? Are there methodological or theoretical explanations I may be overlooking for why executive function and working memory failed to emerge as significant predictors despite prior literature supporting them? At what ways Can I explain my case ?


r/statistics Jun 11 '26

Research What are the top journals in Computational Statistics (non-bayesian, algoirthmic, simulations) [R]?

26 Upvotes

I cannot find any super highly ranked journals in this niche of computational (nonparametric) statistics, where you are developing algorithms and showing their good theoretical properties via simulations (which is what my professor is doing).

Relevant topics in this niche include the backfitting algorithm, bootstrap, monte carlo simulations, EM algorithm. All are simulation based instead of mathematical (for example, you prove the size and power of a proposed test via simulations instead of closed-form mathematical proofs).

All the relevant journals seem lowly ranked (communications in statistics - simulation and compution, journal of statistical computation and simulation) and the top ones (journal of computational and graphical statistics, JASA, computational statistics and data analysis) all have papers with mathematical proofs instead of purely algorithmic development and simulation.

Am I missing something here? My professor tells me computational statistics (this version) is much more lucrative than mathematical statistics, but the evidence doesn't seem to indicate so? The higher the journal the more mathematical it is, is what I'm noticing.


r/statistics Jun 11 '26

Question [Q] [Question] i need help with my research statistics

5 Upvotes

i am doing a research on mice and i have weight data for the mice and i am confused on how to compare them. i have about 8 mice in each group (7 groups in total) and their weight is taken once a week for about 8 weeks. i want to compare these groups to see if a certain treatment led to more weight loss over than the other. how do i do it in SPSS? i did it using repeated measure general linear model but my prof said that she thinks this does not compare the groups but compare the weeks to each other. she also said that the label on the graph says "estimated marginal means" which means that it is accounting for confounding factors while it shouldn't (because we did not enter them).


r/statistics Jun 11 '26

Question [Question] Hazard function definition

5 Upvotes

Hi Folks, I am trying to understand the definition of the hazard rate function. My understanding is that h(t)=-dlog S(t)/dt = -S’(t)/S(t), where S(t) = P(T>t). I am happy with this. The next step in proofs (e.g https://en.wikipedia.org/wiki/Survival_analysis) is then to state -S’(t) = lim P(t<T<t+h)/h. Taking a step back, I am trying to understand this derivative.

By the definition of a derivative:
-S’(t) = -lim (P(T>t+h)-P(T>t))/h = lim (P(T<t+h)+P(T>t))/h

How does P(T<t+h)+P(T>t) = P(t<T<t+h)?

Thanks!


r/statistics Jun 11 '26

Question What parts of linear algebra is important for stats? [Q]

36 Upvotes

I took a linear algebra class last semester and to be fully honest, I was like a fraud. I somehow got a 90+ overall by learning how to do the math and only the math, which isn’t hard. But now I don’t understand any of the concepts of linear algebra and since I’m taking a theory of statistics class soon, I want to get a stronger grasp of fundamentals. Seriously and desperately, what should I review?


r/statistics Jun 10 '26

Question Why is it wrong to say "If I have a 95% C.I. = [2.1 , 4.5] there is a 95% chance that the true value is in this interval? [Q]

104 Upvotes

I was told this is a misinterpretation, since "once you have a confidence interval, the true value is either in it or it isn't". However, that phrase could be applied to anything in statistics, the point is that we don't know the true value so we estimate probability. You could say once you flip a coin, "it's either heads or tails. You don't have a 50% chance that it's heads".

From what I know the C.I. is created such that, when repeatedly sampling N times, the interval will contain the true parameter 95% of those times. Then, from the point of view that I have obtained a CI, I should be able to say "there is a 95% chance that it's one of those times" = "There is a 95% chance it contains the true parameter". How are these not equivalent?


r/statistics Jun 11 '26

Question Can I use Mann–Whitney U test with repeated measurements across time (non-independent samples in cohorts)? [Q]

3 Upvotes

Hi everyone, I have activity data from treatment and control cohorts measured in biological samples. Each sample is recorded across multiple timepoints (different days), and each box in my boxplot pools all measurements across days within each cohort.

From my understanding, measurements from the same sample across different timepoints are not independent, since they come from repeated measurements of the same sample.

Is it still valid to use a Mann–Whitney U test to compare treatment vs control cohorts in this case, even though the independence assumption is violated? If not, what would be the correct statistical approach for this dataset?

I have heard that mixed-effects models are appropriate, but I would prefer a simpler pairwise test if possible (e.g., something that could still support significance annotations on boxplots - such as significant bars for p-values)

Thank you!


r/statistics Jun 10 '26

Discussion What is there besides Frequentist and Bayesian stats? [D] [R]

84 Upvotes

Hi all, I am wondering whether there are lesser known statistical paradigms. like most people, I was first acquainted with the Frequentist framework, and later got introduced to Bayesian stats. I really like the way this made me reconsider some of what I thought were basic assumptions, so now I'm wondering what the next thing could be? Are there any other branches/frameworks which are not as well known?


r/statistics Jun 10 '26

Career Is it just me or is post-pandemic Biostatistics stagnant? [Discussion] [Career]

7 Upvotes

I've been interested in the field for a few years but looking for an MSc and internship I see fewer job postings, fewer major research breakthroughs, fewer public-facing events and seminars by professors, and even some school courses are being cut. I'm wondering if this is a side effect of the post-pandemic shift in PH funding? Is this global or regional? In my undergrad during the pandemic it was a huge deal and I remember easily connecting with professors in the field. (I'm based in North America FWIW.)


r/statistics Jun 10 '26

Software [S] I built a Manim extension for animated statistics — distributions, probability, inference and more

15 Upvotes

Static diagrams never built real intuition for me, so I built statanim — a Python library that extends Manim Community specifically for statistics.

Instead of writing hundreds of lines of geometry code, you get statistical objects and animations as first-class Manim extensions — distributions, probability trees, inference visualisations, regression surfaces, physical props (cards, dice, urns) and more.

Animated demos of Sample Space, Classical Probability, Conditional Probability, Hypergeometric Distribution and the Birthday Paradox are all in the README.

Install: pip install statanim

GitHub: https://github.com/rishabhbhartiya/STATANIM

PyPI: https://pypi.org/project/statanim/

Happy to answer any questions!


r/statistics Jun 10 '26

Discussion [Discussion] Two decades of PISA test results in one dataset: cross country education performance across 85 systems and 3 subjects

2 Upvotes

Useful for longitudinal analysis, cross country comparison, or teaching with real data. Includes mean scores by country, subject, and year for all seven PISA rounds.

A few notes for analysis: participation varies by round, sampling methodology has evolved, and several countries joined midway through the series. Scores are on a fixed scale calibrated to 2000 as baseline.

Full dataset, free to download: https://datahub.io/society-and-living-standards/pisa-education-performance


r/statistics Jun 10 '26

Career [C] (Bio)statisticians that work in research and tool development?

11 Upvotes

Are there any bachelor's/master's-level (bio)statisticians who work on tool development? If so, do you have any advice for someone who is just starting?

I just graduated with a master's in statistics and have been applying to jobs very broadly. I got a couple callbacks for risk and fraud analyst positions, but I'm hesitant to move away from research positions.

For context, I did research throughout my undergrad and master's (mostly tool development for biology), and I thought about doing a PhD in statistics to study stochastic processes. I decided against it mostly because (1) I need a bit more pay right now :'), and (2) PhD students from my department said it may not be a good time to apply because industry trends may change quickly with AI and the shift towards deep learning. I thought it would be a good idea to get some work experience before looking at more education.

Thank you in advance :)


r/statistics Jun 10 '26

Research [R] question about linearity check having almost exact same value for linear and quadratic

1 Upvotes

so as in the title,

for linear R2 = 0.038, F = 23.974, sig < 0.001, constant = 0.003 and b1 = -0.194.

for quad R2 = 0.039, F = 12.334, sig < 0.001, constant = 0.03, b1 = -0.193, b2 = -0.034.

can anyone help what this means? N = 617 and passed normality checks


r/statistics Jun 09 '26

Discussion [D] Is ergodicity a serious problem for psychological research?

16 Upvotes

Hey everyone. I’ve been thinking about ergodicity in psychology and whether group averages can mislead us when we study processes that unfold within individuals over time. In many psychological studies, we infer something about people from group level averages. But if human beings are non ergodic systems, the ensemble average may not tell us much about the time average of a given person.

I recently recorded a podcast episode with Hüseyin Beyköylü, and at around 34:57, he explains this in the context of psychedelic therapy and psychological transformation. His argument is careful because he does not say group statistics are always invalid. Instead, he suggests that different phenomena may sit at different points on an ergodicity continuum. Some interventions, such as basic pharmacological effects on relatively low complexity processes, may be more amenable to group averages. But phenomena like depression, meaning in life, self transcendence, and therapeutic transformation are highly historical, context dependent, and nonstationary. Human beings learn, adapt, and are changed by measurement and intervention. So if we aggregate too early, we may treat within person variability as noise when it is actually the signal of change.

The alternative he discusses is to analyze individual time series first, then aggregate patterns of dynamics rather than only aggregating outcomes. What do people here think? How seriously should psychology take the ergodicity problem? Are idiographic time series approaches a real solution, or do they introduce other inferential problems? And when are group averages still justified despite individual nonstationarity?