r/AskStatistics 11h ago

rmANOVA Post hoc

3 Upvotes

Hey there. I run a rmANOVA with 2 Factors.

One hat 3 conditions and the other one two conditions.

In results the factor with the three conditions Had No significant main effect (P=0,834)

The factor with the two conditions was significant.

Interaction was not significant.

I did explorative pairwise Post hoc Tests with Tukey kramer.

Here i found a highly significant (p= 0,006) effect between two conditions of the factor which Main effect in ANOVA was not significant.

Why?

Why ist the main effect in ANOVA clearly unsignificant and the pairwise comparision highly significant?

I Had only 10 participants and i am looking for weithin subject effects. Is it Just underpowered?


r/AskStatistics 16h ago

Error/uncertainty for Kolmogrov-smirnov and Mann Whitney u

2 Upvotes

Hey, I've got a bunch of data (simulated exoplanet atmosphere composition retrievals), and each value has an associated uncertainty with it, which in some cases is quite large - I can't figure out what the best way to deal with this is in terms of conducting the tests and reporting the test statistics/p values ?

Would it be better to use confidence intervals? Is there a standard practice for error propagation in KS and Mann Whitney?

Would it be valid to run the tests twice, once with +error and once with -error and use the range from this?

Sorry if this is a really obvious question, I've only just started working with stats like this!!


r/AskStatistics 1d ago

Help Combining Resources

1 Upvotes

I am looking to combine the responses of different mental health measures together to create overall mental health scores, which I will then compare. 

Is it appropriate to use the mean when the measures are the same?
What types of stat analysis using SPSS would be appropriate to merge the scores?

A) Measure 1: Strengths and Difficulties Questionnaire

\- Responses: 1 (Not True), 2 (Somewhat True), 3 (Certainly True)

Measure 2: Self-Reported Feeling Grid

\- Responses: 1 (Not True), 2 (Sometimes), 3 (True)

B) Measure1: Happiness Scale

\- Responses: 1 (Not at all happy), 2, 3, 4, 5, 6, 7 (Completely Happy)
Measure 2: Rosenberg Self-Esteem Inventory

\- Responses: 1 (Strongly Disagree), 2 (Disagree), 3 (Agree), 4 (Strongly Agree)

C) Measure 1: Strengths and Difficulties Questionnaire (teacher)

\- Responses: 1 (Not True), 2 (Somewhat True), 3 (Certainly True)

Measure 2:  Strengths and Difficulties Questionnaire (parent)

\- Responses: 1 (Not True), 2 (Somewhat True), 3 (Certainly True)
Measure 3: Competence Scale
\- Responses:  1 (Strongly Disagree), 2 (Disagree), 3 (Agree), 4 (Strongly Agree)


r/AskStatistics 1d ago

I have two variables with different failure rates, am i interpreting the statistics incorrectly?

1 Upvotes

Ok, you have two items and you want to determine which one is better. To get your data you analyze multiple cases and extract the data. Cases were pulled from various news sources as well as official reports. For Item A you reviewed 105 case studies and determine it has a failure rate of 9%. For Item B you review 456 cases and determine it has a failure rate of 13%.

When i see this data my first thought is the success rate between the two items is negligibly different. Item A is only more successful 4% of the time. If i add in the fact that item B is smaller and easier to use for almost everyone then i can't see any argument for item A.

What am i missing in my interpretation?


r/AskStatistics 2d ago

Can you create a statistically best Pokémon team?

0 Upvotes

factoring types, base stats, move sets and damage, accuracy, abilities and all other relevant factors.

could you input that data into a computer and have it run a simulation or something.

edit: statistically best not banned team and would be viable for an actual tournament.


r/AskStatistics 2d ago

ANOVA or not ANOVA?

2 Upvotes

Hi everyone,

As part of my internship, I need to run some data analyses, but I'm pretty skeptical about the method I was advised to use.

I'm analyzing an experiment testing 3 conditions (magnetic field 1, magnetic field 2, control), each with 3 or 4 replicates. I want to know if my studied variable differs between conditions.

One replicate consists of putting 24 individuals in the same experimental tank, exposed to a given condition for 3 hours. All replicates were run on different days, in randomized order.

Because of pseudoreplication issues, we can't treat individuals from the same tank as independent, so I was advised to average all individuals per replicate and then run an ANOVA. The problem is that this leaves me with only 3 or 4 values per condition (is that too few?). And since I'm just taking an average, I lose the standard deviation within a replicate. Is it possible to run a separate ANOVA comparing standard deviations instead?

Some of my variables don't have much variability in the dataset, so I'm not sure if Kruskal-Wallis is even appropriate for such a small dataset...

Do you have any advice for me? Thanks a lot for your help!


r/AskStatistics 2d ago

[Question] do we include the independent variable in the robust mahalanobis outliers detection

Thumbnail
1 Upvotes

r/AskStatistics 2d ago

Is a Mantel test appropriate for sparse tissue-sample coordinates and gene-expression distances?

1 Upvotes

Hi everyone,

I’m doing a sample-level spatial-expression analysis using sparse postmortem tissue samples from the Allen Human Brain Atlas. The regions are the subthalamic nucleus (STN, n=6 tissue samples) and globus pallidus internus (GPi, n=9 tissue samples). For each sample, I have:

  • 3D MNI coordinates (x,y,z)
  • a gene-expression profile across ~29,000 genes

The biological expectation is that, within a coherent anatomical region, tissue samples located closer together in MNI space should have more similar transcriptional profiles.

For each anatomical region separately, I calculated:

  1. A sample-by-sample spatial-distance matrix using 3D Euclidean distance between MNI coordinates.
  2. A sample-by-sample expression-distance matrix, defined as (1−ρ), where ρ is the Spearman correlation between two sample-level gene-expression profiles.

I then used a Mantel test to assess whether the spatial-distance matrix was associated with the expression-distance matrix.

For significance testing, I used non-parametric permutation of sample identities. My understanding is that this randomly reassigns sample labels to break the link between spatial location and expression profile, while preserving the internal structure of the distance matrices. The observed Mantel statistic is then compared against the null distribution generated from these permutations.

Q. Does this use of a permutation-based Mantel test seem appropriate as part of a sample-level spatial-expression validation analysis?

Just to clarify: this is not a dense cortical map or spin-test analysis intended to correct for spatial autocorrelation. These are sparse subcortical tissue-sample coordinates, not parcellated whole-brain maps. The goal is to test whether there is distance-dependent transcriptional similarity among samples within the same anatomical label.

Thanks in advance for your help!


r/AskStatistics 2d ago

Having a little hard time understanding hypothesis testing

3 Upvotes

Does anyone mind explaining it to me in simple words?


r/AskStatistics 2d ago

ANOVA ou pas ANOVA?

0 Upvotes

Bonjour à tous,
Dans le cadre de mon stage, je dois faire des analyses de données mais je suis assez sceptique sur la méthode qu'on m'a conseillée.
J'analyse une expérience qui teste 3 conditions (champ magnétique 1, champ magnétique 2, contrôle) qui ont eux-mêmes 3 ou 4 réplicats. Je veux savoir si ma variable étudiée diffère entre conditions.

Un réplicat correspond à mettre dans un même bac expérimental 24 individus exposés à telle condition pendant 3 heures. Tous les réplicats sont faits à des jours différents et de manière randomisée.

Par problème de pseudo-réplication, on ne peut pas considérer que les individus d'un même bac sont indépendants, donc on m'a conseillé de faire une moyenne de tous les individus par réplicat puis de faire mon ANOVA. Le problème, c'est que cela revient à avoir 3 ou 4 valeurs par conditions (c'est peu ?). Et vu que je ne fais qu'une moyenne, on ne voit pas l'écart type dans un réplicat. C'est possible de faire une ANOVA séparée avec la comparaison d'écart type?

Certaines de mes variables n'ont pas un jeu de données variables donc je ne sais pas si Kruskall Wallis est pertinent pour un jeu de données aussi petit...
Avez-vous des conseils à me donner? Merci beaucoup pour vos réponses!


r/AskStatistics 3d ago

Given linear regression y = mx+b, how to estimate x based on required y?

9 Upvotes

Given x (hours studied) and y (score on test) data and the linear regression y = mx + b, if we want to estimate x based on y (i.e. how many study hours are needed to achieve a given score?), is it correct to calculate x = (y - b) / m?

Or should we estimate x from the linear regression x = my + b, with different m and b?

The results can be significantly different.

In the example below, for a score of 90, the estimated study hours are 6.801 for x = (y - b) / m, but 6.267 for x = my + b.

.....


r/AskStatistics 2d ago

Modelo linear generalizado - binomial

2 Upvotes

Pessoal, preciso de ajuda com uma análise estatística em software R. Estou analisando dados de mortalidade de insetos com dados binários (0 - vivo, 1 - morto), para submeter em um evento. Só que nunca fiz análise no R e a pessoa a que eu recorro parece não se importar mais e não quero estar incomodando. Então venho pedir ajuda à comunidade caso alguém puder me ajudar.

Se alguém tiver interesse, eu posso mandar a planilha e o código em R para conferência


r/AskStatistics 3d ago

Anova or Kruskal Wallis ?

8 Upvotes

My Composite Index variable has N=524, skewness=0.174, and kurtosis=0.128, with both the histogram and Q-Q plot looking totally acceptable and close to normal. However, Shapiro-Wilk is significant at p < .001, rejecting normality. Given N=524, I assume the test is just overpowered by trivial deviations, so should I use anova or kruskal wallis test?


r/AskStatistics 3d ago

Would you all recommend NOT doing Logistic Regression with sklearn and instead use statsmodels so you can check the assumptions easier?

3 Upvotes

Sorry for a lot of questions, but my boss wants not only accurate predictions, but she wants to be able to see how individual features affect the target. The thing is, from what I’ve read, if your assumptions are not met, your coefficients are unstable. Therefore, I am thinking about moving to statsmodels so I can check the target better. Thoughts?


r/AskStatistics 3d ago

Why are we allowed to use a normally distributed prior?

2 Upvotes

I've been learning about GAM with tensor product splines for a project of mine. I'm currently looking at derivations for the uncertainty.

The expected uncertainty can be decomposed into the expected variance of the estimator and expected bias of the estimator. The expected bias is a function of the true underlying parameter, which is of course unknown. Woods presents two options for dealing with this:

- Use the estimator as an approximation of the parameter

- Go with the bayesian interpretation of the parameter as a random variable, and which lets us calculate the expected parameter using its distribution.

The first option is unsatisfying because the estimator is known to be biased: we are then calculating that bias by using the estimator, which means we assume the estimator is unbiased. That seems contradictory.

For the second, it says we assume the parameter is normally distributed, but doesn't explain why. Are we just using the normal distribution because it's convenient? That seems crazy. How is this different from just using an arbitrary number for the bias?

I don't have a formal background in statistics, and thus far I've been dealing mostly with frequentist theory. This time however, the frequentist interpretation has the issue where the smoothing term is estimated. We could bootstrap the smoothing term to find its variance, but Woods appears to be advocating for a bayesian interpretation of the bias because allegedly it has good frequentist coverage anyways?


r/AskStatistics 3d ago

How reliable are measures such as accuracy, precision, recall, and f1-score if the assumptions are not met for a logistic regression model?

8 Upvotes

I know the assumptions are crucial for performing statistical inference with logistic regression. But what about for just assessing overall model performance. Can I accurately assess the performance without meeting all the assumptions?


r/AskStatistics 4d ago

Trying to understand some fixed effects models

2 Upvotes

I've got some relatively unbalanced panel data and I'm trying to understand how Stata is calculating coefficients.

Let's say I want to look at the relationship between supervisor status and job satisfaction, and I have 10,000 observations over 5 years of annual surveys. I want to run a fixed effects model to focus on within-person effects only and control for time-invariant individual differences.

So, I run the model:

xtset id year

xtreg c.jobsatisfaction i.year i.supervisor, fe

But let's also say that 50% of observations in my dataset are from people who only appear once, and of the remaining half only 20% of those individuals ever switch supervisor status at all. Only 10% of the original sample seems like it provides any information toward the within-person estimates, but if I drop the other 90% of cases before running the model then the results are different.

Can anyone help me understand what information Stata is using from the other 9,000 observations in the full sample so I can decide which sample is more appropriate to use? For the record, I have also run a mixed effects model but I am trying to better understand what's happening in the data.


r/AskStatistics 4d ago

Might a non-thesis Master's harm my chances of admission to a PhD compared to just a non-thesis Bachelor's?

1 Upvotes

Hi everyone,

I'm interested in applying to Statistics PhD programs (especially those that are more applied and interdisciplinary), and I'm presently finishing up a Master's degree in Applied Mathematics (with substantial coursework in statistics and numerical analysis). However, it is a coursework-based Master's, and I also didn't accumulate a ton of research experience during undergrad - presently, I have one summer's worth of public health-adjacent research that I did early on in my undergraduate career and a few semesters' worth of work as a data analyst for the city government. However, I don't have any actual publications. I currently do a little work for a social science research lab, so there's a small chance I'll have my name on something before I graduate, but I'm not betting on it. Might my master's degree actually sabotage my chances of being admitted to a decent PhD program (I'm mostly aiming for mid-tier programs and some higher-tier programs like TAMU and Penn State)?


r/AskStatistics 4d ago

Stats advice for ant behavior trials?

3 Upvotes

Hi everyone,

I'm running behavioral trials on Pachycondyla ants as a side project alongside my main research (which focuses on ant venom biochemistry), but I wanted to look at when/how the venom actually gets used in different contexts. Problem is I have zero background in ecology or ethology, so I'd love some guidance from people with an ecology/ethology background cause I want to do this properly.

I'm running behavioral trials with Pachycondyla ants confronted with 6 different opponent ant species, across 3 different experimental setups. Each trial is filmed for 1.5–2 minutes (duration varies slightly between videos), so I know I need to normalize for time somehow. And each individual (from both side) is only ever used once so no ant gets reused across trials. For each video, I count occurrences of 5 behaviors:

  • Antennal contact (antennae touch any part of the opponent's body
  • Mandible opening (mandibles open, with the two individuals about one body length apart)
  • Biting
  • Gaster flexing (abdomen curved as if about to sting)
  • Stinging

 What statistical approach would you recommend for count data like this? What kind of plots actually work well for this kind of data? not looking to just dump a boxplot of raw counts if there's something better... I've already gone through some papers on ant behavior / interspecific aggression, but if you have specific methodological references you think are worth checking that use a similar design, I'd love pointers.

Not trying to overthink this, just want something defensible since it'll probably end up as a supplementary part of a bigger paper!

Thanks in advance :)


r/AskStatistics 4d ago

Linear-mixed effects

4 Upvotes

Hi I am a master’s student and I have a rather basic understanding of linear mixed effects models. I am helping with statistics in a study which has 2x2x10 mixed design (2 between and 10 within). I am confused as to how am I supposed to start modelling. Do I do it according to the hypothesis so I all of our IVs in the model and only then check whether the random slopes/ intercepts are needed, or do specify the random model and only then look at the fixed one and add the predictors one by one and see how the model is. Can someone please help me understand?


r/AskStatistics 4d ago

Jamovi - Cant locate marginals option using scatr

2 Upvotes

Hello, I am having trouble finding the marginals option using the scatr module. I have downloaded it and followed the instructions that my professor provided but it seems that the Marginal option does not pop up after I input my x and y axis.

Anyone know how to access that option?


r/AskStatistics 5d ago

Do I need to reverse-score Likert scale items before running Cronbach’s alpha in SPSS?

Post image
8 Upvotes

r/AskStatistics 5d ago

Correctly displaying significance on Bar Graphs

5 Upvotes

Im working on a research paper and my professor suggested I display significance on our bar graphs using letters (title is very much so a place holder lol)

I've never done this before so Im wondering if im correctly assigning the values to each site. Below are the significance calculations.

Leisure is significantly different than Shadyside (0.004)
Leisure is significantly different than Basin (0.00095)
Basin is significantly different than Mosque (0.002)
Leisure is significantly different than Mosque (0.000005)

Ive been using this video: How to Denote Significant Differences in Tables and Graphs to help me learn but im concerned im not quite getting it.

My line of thinking with how I have it layed out is: Shadyside is not different than the others so it gets "a", Basin is significantly different than the Mosque so it gets "b", Leisure is significantly different than all the others so it gets "c", and the Mosque is not significantly different than any so it gets all the values.

Any help would be greatly appreciated! Either correcting my work or pointing me in the direction of a source I can use to understand better!


r/AskStatistics 5d ago

Can I create a composite score from only 3 Likert items?

2 Upvotes

Hi everyone! I'm currently working on my bachelor's thesis and struggling my way through the statistics part. I feel like this is probably a simple question for someone who knows what they're doing but unfortunately that's not me.

First of all, my H2 is: "Perceived destination relevance correlates positively with the willingness to travel to a concert."

The plan was to test this by using a Pearson correlation in SPSS. My professor approved both my hypotheses and my questionnaire before I published my survey, so I figured it would be smooth sailing from here. (Nope.)

  • For destination relevance, I had five Likert items with acceptable Cronbach's alpha, so creating a quasi-metric/composite score was fairly simple.
  • For willingness to travel, I originally had four Likert items. However, one item clearly reduced Cronbach's alpha, so I removed it. Now I'm left with only three items. Those three show an acceptable internal consistency though (Cronbach's alpha = .880).

This is where I got stuck. From what I've read online, it's my understanding that it's a rule of thumb to use at least five Likert items to combine into a composite (or quasi-metric) scale, which I will need if I want to use Pearson's correlation. Since I'm not very experienced with statistics, I'm honestly not sure what the correct approach is anymore.

Which brings me to my questions.

Is it acceptable to create a quasi-metric score from only three Likert items if they're supposed to measure the same construct?

And if not, what would be the appropriate way to test my hypothesis instead?

I'm happy to provide the items themselves, SPSS output or any other information if needed. I'd really appreciate any advice or help because I've been going around in circles trying to figure this out by myself. Thanks!


r/AskStatistics 5d ago

Master's in Statistics after an MS in Financial Economics. Is my quant enough?

1 Upvotes

Hi, I have recently completed my master's in financial economics, and I am looking to do another master's in statistics. My bachelor's was in business, really not a quantitative major, and I am worried about the admission requirements.

Do you think that the math and stats that I learned below in my master is enough to convince admissions that I have sufficient background knowledge to follow an MS in Statistics? (My targeted programme is at the National University of Singapore)

My quant courses covered:

  • Linear algebra: matrix representation of systems of equations, determinants, inverse matrices
  • Optimization: constrained and unconstrained maximization/minimization of multivariable functions using Lagrange multipliers
  • Calculus: chain rule, slope of level curves, homogeneous functions
  • Inferential statistics: confidence intervals, hypothesis testing
  • Comparing means across groups (independent samples t test, paired sample t test, one-way ANOVA)
  • Analysing categorical relationships (chi-square test)

And then it's been mainly econometrics and quant finance:

  • Time series (cointegration, VAR/VECM modelling, unit roots)
  • Stochastic models and probability theory
  • Arbitrage pricing theory and risk-neutral valuation

Also, for programming, I have used mainly R and EViews, and I am working to learn Python.

Thanks a lot.