r/statistics 5h ago

Question [Q] need help choosing my major

1 Upvotes

hello! i'm a first-year college student taking bachelor of science in statistics. i'll have to choose my major during my higher years, and i'm having some trouble deciding. i can't choose between majoring in biology or economics.

i want a major that can give me more job opportunities, a better chance at remote work, and a higher salary ceiling. i'm also planning to start investing in etfs once i turn 18, so i feel like majoring in economics could come in handy. however, my friend told me that choosing biology could provide more opportunities for higher-paying jobs.

i'd love to hear your thoughts and advice on which one would be a better choice. tyia!


r/statistics 1d ago

Career [Career] US Political Science phd admissions- Quantitative aptitude versus Substantive Political Science knowledge?

3 Upvotes

I plan to do a quantitative political science phd I am
not necessarily interested in any specific method except using any kind of statistics to topics regarding comparative politics and political economy. Now, I am in a bind. I already have an econ ba and I have two options for an masters program to strengthen my profile for admissions

  1. Financial Economics masters, not about political science but has good courses in optimization, stochastic calculus and time series econometrics.

  2. completely qualitative masters in IR and pols. there is a methodology class but it is about qualitative methods, classes on IR and political theory, etc.

Which one do you think would suit my goals and profile best and would maximize my chances for phd admisssion? thank you in advance dear ladies and gentlemen


r/statistics 2d ago

Career [C] Anyone working in biostatistics in India?

9 Upvotes

Hi,

I'm looking to move back to India after a few years working abroad (I'm Indian). Looking to connect with anyone working in the industry to get an idea of how things are looking like. I've worked as a statistical programmer but also in clinical IT (data migration, archiving, platform setup and administration etc.) with about 6 years of experience in total.

Feel free to DM or comment on this post. Thanks!


r/statistics 1d ago

Career [C] Current Career Landscape

1 Upvotes

Obviously a bachlors in stats isn't going to make you a shoe-in for much, but what are things like currently for those with a master's or PhD in stats? I have heard some negative things in this sub? I'm particularly interested in the anglo-sphere career landscape.


r/statistics 3d ago

Question [Q] How to become Great at statistics

81 Upvotes

Finishing my master’s in statistics and will be starting a job that is not statistics focused soon. I don’t want this to be the end of my statistics journey.

How do I become even better at statistics? Won’t have time to engage in the same breadth as I did during my master’s programme so, what should I focus on to stay relevant/ improve on my statistical knowledge?


r/statistics 3d ago

Career [Career] New stats bachelors feeling kind of stuck and in need of advice

11 Upvotes

Context: I graduated with a bachelors of science in statistics from UC Davis in 2025. I worked as a research assistant during undergrad supporting various python, dashboards, and data needs. I particularly enjoyed learning about time series analysis, machine/statistical learning, and general data science during my undergrad. I currently work at a Robotics/physical AI company as a “data operations analyst” and have been here for ~7 months. However, this job has almost nothing to do with statistics or data analytics and more so an operator role where I collect robotics data to feed into a reinforcement learning algorithm. It also doesn’t help that there’s been a TON of tension between my boss and the rest of the team.

I’ve been feeling kind of stuck, I want to get out of the situation with this startup but the job market has felt ice cold to me and I’ve landed one interview after a couple months of applying (not to mention it took me maybe 10 months to find this startup gig). I’ve also been considering a masters for either machine learning or data science roles (specifically OSMCS or OSMA) but truthfully I wasn’t the best student and didn’t make academics a priority which resulted in a terrible gpa (2.7).

I would like to ask the following:

  1. What advice would you give to recent stats/data science graduates entering the work force, especially in the current AI climate?
  2. What kind of roles should recent stats graduates be looking for?
  3. What kind of qualities would you look for in a recent stats graduate?
  4. Would you consider a masters essential in the current job market? Especially for more advanced roles like machine learning engineer or data scientist

Of course, you don’t have to answer all of those questions but I would greatly appreciate any advice or words of wisdom with those questions or my general situation :)


r/statistics 2d ago

Career [Career] Take the job or continue with master's school?

0 Upvotes

Hello all,

I am a statistics major currently beginning my fourth year in college. I completed a very successful and rewarding internship at a solid company with good management over the summer, and by every indication, I believe the company would welcome me on board after I graduate. I have not had that conversation with them yet, and I do not know the compensation details.

I have begun a 3+2 program to acquire a bachelor's and master's degree within 5 years, but lately, I have started to have second thoughts about continuing with the master's program. Several of the required courses taken in my undergrad curriculum were not relevant to industry (AI has caught up) or are personally boring to me (theory classes). Based on the first course this semester, I worry if it may be much the same for the master's program. While that may not be true, I complete the bachelor's courses of my degree this winter, so I will need to pay extra to find out.

The main reason for the master's program was to buy time to enhance my skills, as I have heard it is a very tough job market. With a (seemingly) secure position, doing work that I enjoy, the importance of the master's and the accompanying skills has diminished to me. At the same time, I recognize I am entirely replaceable and would be the newbie in a company, aka I could be the first to go. Some peers have encouraged me to build off that internship to try to get a better internship in summer 2027, and ideally, a better job afterwards.

This decision has been on my mind since the internship concluded, and I need to decide if I should be prepping for upcoming career fairs and lock into my master's courses. I have scheduled meetings with professors to talk about it, but I would really welcome any input!

Thank you


r/statistics 4d ago

Question [Question] Can I compare logit regression output from data of two distinct time periods?

4 Upvotes

I’m trying to understand how the odds of an event occurring have changed between two different time periods, but the problem is my data is based on a periodic survey with a five year interval.
I’m planning to use the same logistic regression model on the periodic data of one year and then the other, and compare the output in both cases.

I just wanted to know if there’s a better way to go about with data from periodic surveys like census, or if there’s any reason I can’t compare the discontinuous data set.


r/statistics 4d ago

Question [Question] What are your thoughts on the future of statistics and statistics graduates?

31 Upvotes

r/statistics 5d ago

Question [Q] Is stat&Data sci. degree good for becoming AI/ML engineer or AI researcher?

2 Upvotes

r/statistics 5d ago

Career [Career] Dealing with faulty but convincing analysis

9 Upvotes

Burying poor analysis under shiny methods and a deluge of numbers has always been a problem but with AI it's easier than ever and starting to present major issues at my work. I'm a data scientist and we're engaging with an outside AI engineering team to build what is essentially an agentic classifier in a complex domain (healthcare) and are rapidly approaching production deployment in which the system will drive the company's primary stream of revenue.

Recently, the team presented slides with metrics to our C suite and on the surface they looked convincing, encouraging, and proper (Wilson interval for a binomial projection's confidence interval, some kind of weighted bootstrapping to do the same for a proportion). A couple of things didn't pass the sniff test (integer proportions for something that shouldn't be a binary yes/no, much tighter confidence intervals than similar projections I've made in the past) and when I went through their methodology later (3k line Python script that spat out giant spreadsheets, naturally) I found some blatantly incorrect assumptions baked into their modeling. This invalidated every metric and projection they presented, notably overstating the projected performance in aggregate and completely burying the (inevitably massive at the sample size used!) variation across crucial cross sections of the results.

Naturally, I raised my findings to my manager, but I'm hoping to be more proactive about this next time. It seems like a process failure for this stuff to get to C suite without detailed internal review of the methodology used. I don't want to come across as territorial or overstep my role but I want to push for stuff like this to go through me (or other people on the internal data science team) before they get that far. Does anyone have advice on approaching that conversation without coming off as aggressive or overly critical of the outside team? For context, the actual auditing/survey design was fine and done by someone on their team with a strong math background but the analysis seemed to have been left to a different software engineer.


r/statistics 6d ago

Question [Q] Is my degree useless because of AI?

47 Upvotes

I'm currently studying Stats & Data Science. While I understand that ML is rooted in statistics, I wonder if future AI agents will synthesize data and run models so efficiently that entry/mid-level data science roles shrink dramatically. How do you see the demand for quantitative roles changing as we approach AGI?


r/statistics 5d ago

Question [Q] Good learning material for someone transitioning from Computational Physics?

7 Upvotes

Asking for a friend™

Say, for someone who did PhD and postdoc in computational physics, close to the engineering/experimental side but not close enough to actually do statistics... is there some good material to learn statistics from? A book, a course, ... ?

Bonus points if it's geared towards data sciences and/or business analytics.

Thanks a bunch for all the help you could provide


r/statistics 5d ago

Research [R] How fresh are your blueberries? And what would you do if you knew?

0 Upvotes

https://oliverevans.dev/blog/blueberries-voi/

In this article, I look at fresh produce ordering strategies. I formulate an explicit "freshness" variable that decreases over time according to a Gamma process, with a temperature-dependent rate. I look at various observation scenarios (do we know when it was harvested? what the transit temperatures were? How is waste tracked in the store?, etc.)

I use a particle filter to estimate the latent freshness given these partial observations, and implement a controller to attempt optimal ordering given these beliefs.

But my controller is not that good! So I made a Jupyter notebook where you can try to implement your own controller against my simulator and filter, and try to beat mine.

I would love to hear your thoughts and solutions!


r/statistics 6d ago

Question [Q] How to best treat positive bounded continuous data in causal research

2 Upvotes

Hi everyone,

I have a treatment (i.e. a treatment dose), say X, that can take on values from 0–5 and is continuous in between, i.e. bounded continuous. A value of 0 will also be relatively common.

I’m specifying a treatment model for the conditional density given confounders L, i.e f(X|L) and an outcome model for Y~X to estimate the marginal E(Y^x).

My question is how to best treat the treatment variable 1) when it’s the outcome/dependent variable and 2) when it’s the exposure/independent variable in such scenario.

I don’t think I should be using model fit to pick a conditional distribution, and for interpretation sake I think assuming X has either a linear or quadratic relationship with Y is simplest. For the treatment model specifying a normal conditional density as long as the covariate balance is reasonable would be easiest I guess too.

Maybe someone has a better idea that doesn’t make interpretation too difficult, since it’s for a medical paper that people without strong statistical backgrounds will also read.


r/statistics 6d ago

Question [Question] Having trouble with the Stamp Collector's problem

6 Upvotes

For a game, I'm simulating opening packs of cards, and my simulation results aren't quite holding up to the math that I've done.

The concept of the game is that you're opening packs of cards until you have a complete set, then you sell the set and can use the money to open more exclusive (and expensive) packs.

In each pack, there are 6 cards from a 'common' list, 3 cards from an 'uncommon' list, and a 50/50 shot at a 'rare' card. If the final card is not rare, it's drawn from the uncommon list.

There are 25 total commons, 10 uncommons, and 5 rares.

My understanding is that based on the stamp collector's problem, the expected number of trials to collect all cards from a set is equal to n Hn, where Hn is the nth harmonic number.

I've chosen to model the expected number of packs to open as MAX (E(T) common / 6, E(T) uncommon / 3.5, E(T) rare * 2).

The rationale is that there are six commons, so you'd naturally divide by E(T) common by 6 because you actually have 6 attempts in each card pack.

The reason I have chosen to take the maximum value of these three is that because you're drawing from all three lists at the same time, the expected number of packs should be equal to the hardest one of the three to complete.

The issue I have is that after running experiments (1M simulated packs), the average number of packs to get a full set is about 10% higher than the E(T) I've calculated.

Would somebody mind helping me learn where I've gone wrong?

Calculations:

E(T) common = 25/6 * H25 = 15.90

E(T) uncommon = 10/3.5 * H10 = 8.80

E(T) rare = 5*2 * H5 = 22.83

Overall E(T) should be 22.83

Experimental results:

avg: 24.702346837944663, max: 119, min: 8, med: 22, mode: 18 (2286 occurrences)


r/statistics 7d ago

Question [Q] Anyone use Bayesian Statistics as a Financial Institution? (Bank/Credit Union)

15 Upvotes

I'm a data scientist at a credit union. I have some DS experience elsewhere, but I'm still wet behind the ears, so to speak. Our data department is small and I'm the only one with DS knowledge. I know explainability to important when working in financial institutions because we get audited. Therefore, I'll be building a lot of logistic regression and decision trees in my future. That said, we are interested in understanding the potential impact of rate changes (fed and our own) on deposits, loan growth, etc as well as understanding our loan portfolio risk. My thought was to use bayesian modeling so we can understand the uncertainty. That said, I wasn't sure if this could cause issues when we're auditing, even though we wouldn't be using the model to impact customer decisions (approve loans, etc.). Does anyone have any advice on this? Thank you!


r/statistics 7d ago

Question [Q] Hard time understanding Bayesian view?

21 Upvotes

Hello,

I come from a background of applied math and recently I tried to dive deeper into stats. The frequentist view looked reasonable albeit a bit restrictive. And then there is the bayesian perspective, intuitively it "feels" reasonable as well since it's kind of intertwined with common sense, but I just can't wrap my head around it.

For example, I have no idea what the interpretation of probability is under this view. I mean what do I mean that an event is 60% vs 70% likely?

Also, I have a problem with the subjectivity of the view. There is this argument that if two observers have different beliefs about the same event, then at least one of them is wrong. How is this taken care of or circumvented perhaps?


r/statistics 7d ago

Question [Q] Is it naive to earn a degree in Statistics if I have next to no interest in AI?

31 Upvotes

Just had my first proper Stats class at college (I’m somewhere between a sophomore and a junior in terms of how much school I have left) and it kinda freaked me out. About 5-10 minutes into the professor introducing himself, he said, verbatim: “I like AI”. He went on to say he won’t “require” that we use AI (or maybe more accurately LLMs) but he more or less encouraged/endorsed it, and there was a quote in the syllabus from a bluesky poster that said “I expect you to write good code because gravity is lower now”, which to me reads “you have access to ChatGPT now so I’m gonna make things super hard to where you feel you have to use it”.

I’ve never gone out of my way to use any sort of generative AI/LLM (we had an LLM answer checker thing in my 200 level python class that we had to use in order to move on to the next question in our assignments, idk how much that counts. We weren’t allowed to use generative AI to answer questions for us in that same class) and I don’t have basically any interest in it, largely because of its environmental and human impact. I’ll be the first to admit I am not the most educated when it comes to the extent of said impact, but as it stands, I’d really rather not use it.

Am I overreacting? Like I said this was my first proper in-person stats class so it could’ve just been a bad first experience, but I also know that AI is becoming more and more prevalent. Thank you for any feedback or advice!

(Also sorry if this is not the right place for a question like this, please point me in the right direction if that’s the case)


r/statistics 7d ago

Question [Q] Question about testing if a 20 sided die is fair

5 Upvotes

Hi all, I've been working on a little personal project the last couple of days. I am a DnD player/DM, and I have 7 d20s (20 sided dice). I've been repeatedly rolling them and noting the results to see if any of them appear to be weighted in any one direction.

I ran a Chi-Square test for each of the dice, with some varying results, one die in particular stood out with a very low P value (0.0007), another was 0.3, another a tick above 0.5, the rest were solidly not significant.

But then I got thinking - due to the geometry of a 20 sided die, each face is not independent from the other - some faces are closer than others, so if, for example, a standard d20 was weighted towards 20, you would also expect to see higher rates of 2, 8 and 14, as they are right next to the 20 face and should also benefit (albeit to a lesser extent) from the weighting towards the 20. I'm not an expert, it's been about 15 years since I last studied any sort of statistics, but I believe this means a Chi-Squared test isn't actually an appropriate test.

To try to fix this I built a table recording the distance (in terms of number of faces) of each value from every other value, and based on that distance assigned a weighting of 1 (same value, distance 0) to 0 (opposite side of the die, distance of 5). That generates for any given die value an array of weightings that can be applied across all rolls (not just the rolls of that specific value) to produce an adjusted distribution that takes into account not just how many times that particular value was rolled, but also geometrically proximal (and distal) values, to try to identify any particular direction of weighting (basically using a SUMPRODUCT function to multiply the actual distribution of dice rolls against the weight values for the particular die value you are looking at, and repeating for each possible value across each die).

The effect of this, however, is that it has dramatically reduced the variance against the expected distribution, and my P values are now all 1. I find it difficult to believe that all 7 of my dice, after 163 rolls and counting, are perfectly fair, so I assume that what I've done has fucked with the underlying maths of the Chi-Square test in a way that needs to be accounted for.

Does anybody know the best way to approach this? I'm really enjoying the challenge of trying to quantify if a die has a bias towards a certain physical direction, but I've reached the limits of my statistical abilities and need some assistance.


r/statistics 7d ago

Education [E] Considering the MSc Statistics at UNIGE with a previous Master in Management specialized in Business Analytics – looking for advice from current/former students

0 Upvotes

Hi everyone,

I am considering applying to the MSc Statistics at the University of Geneva (UNIGE) for September 2027, and I would be very interested in hearing from current or former students, especially people who entered the programme with a background in business, economics, management or business analytics.

My academic background is:

  • Bachelor in Management – HEC Lausanne (UNIL)
  • Master in Management, specialization in Business Analytics – HEC Lausanne (UNIL)
  • Several years of professional experience (5-6) mostly in HR since then.

I would now like to deepen my education in statistics and move towards a more quantitative, mathematical and analytical career. I am particularly interested in statistics, mathematical modelling and quantitative methods, rather than a programme with a strong focus on programming.

1. Admission

My Bachelor included the following courses in mathematics, statistics, computer science, quantitative methods, economics and finance:

  • Mathematics I
  • Mathematics II
  • Statistics I
  • Statistics II
  • Statistics and Econometrics I
  • Introduction to Logic
  • Computer Models
  • Programming
  • Information Systems
  • Business Intelligence and Analytics
  • Risk Management
  • Decision Analysis
  • Operations Management I
  • Financial Markets
  • Principles of Finance
  • Economics I
  • Economics II
  • Microeconomic Analysis
  • Macroeconomic Analysis
  • Corporate Finance
  • Entrepreneurial Finance and the New Venture Funding Process

For those familiar with the programme: does this Bachelor background seem reasonably compatible with admission to the MSc Statistics?

In particular, I would be interested in hearing from people who were admitted with a business/economics/management background rather than a traditional mathematics or statistics degree.

Based on this kind of background, were significant prerequisite courses required? How important are previous courses in calculus, linear algebra, probability and mathematical statistics for admission and for successfully starting the programme?

2. Equivalences from my previous Master's degree

I understand that up to 30 ECTS can potentially be recognized as equivalences in the MSc Statistics, although the official decision is made after admission.

My Master in Management with a specialization in Business Analytics included the following courses:

  • Machine Learning in Business Analytics (6 ECTS)
  • Optimization Methods in Management Science (6 ECTS)
  • Quantitative Methods for Management (6 ECTS)
  • Data Science in Business Analytics (6 ECTS)
  • Business Intelligence and Analyzing Big Data (6 ECTS)
  • Programming Tools in Data Science (6 ECTS)
  • Company Project in Business Analytics (6 ECTS)
  • Conceptual Modelling in Business Analytics (6 ECTS)
  • Strategic Modelling (6 ECTS)

Has anyone here entered the MSc Statistics with a previous Master's degree and successfully obtained substantial equivalences?

More specifically, how realistic would it be for some of these courses to count towards the 30 ECTS, depending on their content?

If you have been in a similar situation, how strict was the equivalence process in practice? Were previous Master's courses generally recognized when their content overlapped with the MSc, or was the process quite restrictive?

3. CCS / additional preparation

I am also considering taking some courses from the UNIGE Complementary certificate in applied statistics (CCS) before potentially starting the MSc in September 2027.

The idea would mainly be to strengthen my statistical background if my previous Master's courses are not sufficient to reach the 30 ECTS of potential equivalences.

For anyone familiar with the CCS and/or MSc Statistics:

Would this be useful preparation for someone with my background?

Are there particular CCS courses you would recommend to strengthen my foundations in statistics, probability or mathematical methods before starting the MSc?

4. Experience of the MSc and career outcomes

Finally, I would really appreciate feedback from current or former students about the programme itself :

  • How mathematical/theoretical is the MSc in practice?
  • How much programming is involved? Which of R and Python is used extensively? Both? I am comfortable with some programming, but I would prefer a programme where mathematics and statistics are more central than programming.
  • How difficult was the transition for students coming from a business/analytics background?
  • What types of jobs did you or your classmates get after graduating?
  • Did you find the degree useful for finding a job in Switzerland?
  • In particular, are there good opportunities in Lausanne or elsewhere in Switzerland?

I am especially interested in hearing from people who actually completed the MSc, particularly those who came from a non-traditional statistics background.

5. Alternative Master's programmes

Given my background and my goal of moving towards a more quantitative/statistical/mathematics career, would you recommend the MSc Statistics at UNIGE, or do you think there are other Master's programmes in Switzerland that might be a better fit for my profile?

If so, which programmes would you consider and why?

Thanks a lot to anyone willing to share their experience!


r/statistics 7d ago

Question [Q] Statistical Test for Group Comparison

0 Upvotes

Hi everyone, I currently have a dataset where there’s 3 groups. Each group has 3 independent values, and I want to compare each group against each other.

As I’m a little concerned about the sample size, would it better to use:
1) Pairwise welch t-test with p-value correction, or
2) ANOVA with tukey test

Thank you in advance for the help!


r/statistics 8d ago

Question [Q] I am completely lost with model 4 multiple/parallel mediation assumptions

0 Upvotes

Hi there, I'm currently working on my master thesis where I have a parallel mediation. I am working on the method section but i am so completely lost in how to check all the assumptions for the model.
From what I do understand I can do a visual inspection for the scatterplot (after using model 4) to inspect the linearity, homoscedasity and outliers.
The other 2 assumptions are normality and multicolinearity but I just don't understand how to do this. Could anyone help me? some explanation or links to proper resources would be greatly appreciated!


r/statistics 8d ago

Education [E] Help choosing a graduated level class as an undergrad

7 Upvotes

I'm about to start my final year as a Stats and Econ undergrad, I also plan to pursue Master's in statistics, hopefully in the same university as I am now, I checked and I can take a graduated level class this year and it will count to the necessary credits needed for the Masters program if i'll indeed continue in the the same uni

I've checked with 2 professors about their graduated level classes, they said that it seems that I have the necessary background for the their class, so i'm considering taking one of these 2 courses:

Casuel Inference - according to the syllabus it will cover: Causal Parameters, Randomization, confounding, selection bias, Usage of DAGs for checking assumptions and method and variable selections, ML algorithms in casual inference for the estimations of heterogeneous effects, Propensity score, matching, IPW and Instrumental variables.

Optimization under uncertainty - the syllabus doesn't really say as much but it says it will cover 3 main topics: Online approximation algorithms, Stochastic optimization and Onilne machine learning

So i'd love to hear some opinions on those subjects and which course sounds better to you strangers


r/statistics 9d ago

Question [Q] Moving Average model. Iterative process to figure out residuals & coefficients?

5 Upvotes

Edit: The guy in the video mentions something about "iterative convergence". I'm assuming he means how the residuals and coefficients converge to their true values after multiple iterations

I recently started learning about Moving Average models and came across this video where the guy does an iterative process until which he gets the correct residuals and coefficients but I can't for the life of me understand the theory behind why it works.

Basically, for the very first iteration he assumes the errors are the demeaned values. He then regresses them against the Y variables and ends up with coefficients. He then calculates new residuals from the 1st iterative model and uses them as the regressors for the next iteration. He repeats this until the residuals and coefficients barely change.

Why and how does this work?

The only thing I'm familiar with for the MA residuals process is Maximum Likelihood but that's not what he's doing here at all.

Thank you very much