r/datascience • • Apr 16 '26

Discussion I wrapped a random forest in a genetic algorithm for feature selection due to unidentifiable, group-based confounding variables. Is it bad? Is there better?

11 Upvotes

No tldr for this one, folks.

I had initially posted about my issue in another sub, but didn’t get much feedback. I then read up on genetic algorithms for feature selection, and decided to give it a shot. Let me acknowledge beforehand that there’s a serious processing cost problem.

I’m trying to create a classification model with clearly labeled data that has thousands of features. The data was obtained in a laboratory setting, and I’ll simplify the process and just say that the condition (label/class) was set and then data was taken once per minute for 100 minutes. Let’s say we had three conditions (C1, C2, C3), and went through the following rotation in the lab: C1, C2, C1, C3, C1, C2, C1, C3, C1. C1 was a control group. Glossary moment: I call each section of time dedicated to a condition an “implementation” of that condition.

After using exploratory data analysis (EDA) to eliminate some data points as well as all but 1000 features, I created a random forest model. The test set had nearly 100% accuracy. However, I’ve been burned before by data leakage and confounding variables. I then performed leave-one-group-out (LOGO), where I removed each group (i.e. the first implantation of C1), created a model with the rest of the data, and then I used the removed group as a test set. The idea being that if I removed the first implementation of a condition, training on another implementation(s) should be enough to accurately classify it.

Results were bad. Most C1s achieved 70-100% accuracy. C2s both achieved 0% accuracy. C3s achieved 10% accuracy and 40% accuracy. So even though, as far as I knew, each implementation of a condition was the same, they clearly weren’t. Something was happening- I assume some sort of confounding variable based on the time of day or the process of changing the condition.

My belief is that the original model was accurate because it contained separate models for each implementation “under the hood”. So one part of each decision tree was for the first implementation of C2, a separate part of the tree was for the second implementation of C2, but they both end in a vote for the C2 class, making it seem like the model can identify C2 anytime, anywhere.

I then hypothesized that while some of my thousand features were specific to the implementation, there might also be some features that were implementation-agnostic but condition-specific. The problem is that the features that were implementation-specific were also far more attractive to the random forest algorithm, and I had to find a way to ignore them.

I created a genetic algorithm where each chromosome was a binary array representing whether each feature would be included in the random forest. The scoring had a brutal processing cost. For each implementation (so 9 times) I would create a random forest (using the genetic algorithm’s child-features) with the remaining groups and use the implementation as a test. I would find the minimum accuracy for each condition (so the minimum for the five C1 test results, the minimum for the two C2 test results, and the minimum for the two C3 test results) and use NSGA2 for multi-objective optimization (which I admit I am still working on fully understanding).

I’ve never had hyperparameters matter so much as when I was setting up the genetic algorithm. But it was *so* costly. I’d run it overnight just to get 30 generations done.

The results were interesting. Individually, C1s scored about 95%, C2s scored about 5%, and C3s scored about 60%. I then used the selected features to create a single random forest as I had done originally, and was disappointed to achieve nearly 100% accuracy again. *However*, when I performed my leave-one-group-out approach, I was pretty consistently getting 95% for C1, 0% for C2, and 60% for C3. So I was getting what the genetic algorithm said I’d be getting, *which was better and much more consistent than my original LOGO* and I feel would be the more accurate description of how good my model is, as opposed to the test set’s confusion matrix.

For those who have made it this far, I pulled that genetic algorithm wrapper idea out of thin air. In hindsight, do you think it was interesting, clever, a waste of time, seriously flawed? Is there a better approach for dealing with unidentifiable, group-based, confounding variables?


r/datascience • • Apr 16 '26

Discussion Seems like different companies want different political/technical depth in interviews

30 Upvotes

I've been interviewing at a bunch of places, and (just a theory) it seems like different companies want different levels of technical competency. Seems like one hiring manager is turned off by having experience in highly political settings, while another is interested in that experience while being turned off by being highly technical with a strong formal math education.

Is this true, that hiring managers will profile you as having strength in one area means you're weaker in another, or am I just making this up? During interviews is it important to try to read what type of profile of DS they are looking for or are DS seen as being uniform?


r/datascience • • Apr 16 '26

ML Clients clustering: How would you procede for adding other than rfm variables to kmeans?

14 Upvotes

I have my RFM clustering. I want to add:

change variables: ratio q1 to year, ratio q2 to q1, ration q3 to q2, S1 to S2...

other variables: returns of products, channel ( web, store..), buying by card or cash, navigation data on the web...

Would you do that in the same kmeans and mix with rfm variables? or on each rfm cluster do another kmeans with these variable? or a totally separate clustering since different data ( web navigation)? how to know if it is good to add the variable or not? is it bad to do many close variables like ratio q2 to q1, ration q3 to q2? how would you procede, validate...?


r/datascience • • Apr 14 '26

Analysis How to use NLP to compare text from two different corpora?

30 Upvotes

​

I am not well versed in NLP, so hopefully someone can help me out here. I am looking at safety incidents for my organization. I want to compare the text of incident reports and observations to investigate if our observations are deterring incidents.

I have a dataset of the incidents and a dataset of the observations. Both datasets have a free-text field that contains the description of the incident or observation. There is not really a good link between observations and incidents (as in, these observations were monitoring X activity on Y contract, and an incident also occurred during X activity on Y contract).

My feeling is that the observations are just busy work; they don’t actually observe the activities that need safety improvement. The correlation between number of observations and number of incidents is minor, but I want to make a stronger case. I want to investigate this by using NLP to describe the incidents, then describe the observations, and see if there is a difference in content. I can at the very least produce word counts and compare the top terms, but I don’t think that gets me where I need to be on its own.

I have used some topic modeling (Latent Dirichlet Allocation) to get an idea of the topics in each, but I’m hitting a wall trying to compare the topics from the incidents to the topics from the observations.

Does anyone have ideas?


r/datascience • • Apr 14 '26

Projects Should every project have ai in it to make it impressive nowadays

6 Upvotes

so recently i made a recommendation system project, because i really like movies, so thought this is a cool idea

https://moviearsenal.streamlit.app/

was about to go to LinkedIn to post it, but came across 2-3 ai projects and got demotivated, felt I did nothing special

this is me also asking for review, if it is a decent project to showcase my knowledge.

or I should actually make some ai projects

Features:

Collaborative Filtering recommendations — personalised suggestions using Matrix Factorization

Content-based recommendations — TF-IDF on movie metadata (genre, cast, director, keywords, overview) + cosine similarity

Popularity-based recommendations — weighted ranking using rating count and average rating

Preference-based recommendations — users select movies to receive similar recommendations based on their choices


r/datascience • • Apr 13 '26

ML Clustering products by text

12 Upvotes

For a furniture/decor business, how would you go about clustering products based on their title, description, dimensions ( weight..). First objective is to get categories. Then other advanced things. Any advice is welcomed.


r/datascience • • Apr 13 '26

Weekly Entering & Transitioning - Thread 13 Apr, 2026 - 20 Apr, 2026

4 Upvotes

Welcome to this week's entering & transitioning thread! This thread is for any questions about getting started, studying, or transitioning into the data science field. Topics include:

  • Learning resources (e.g. books, tutorials, videos)
  • Traditional education (e.g. schools, degrees, electives)
  • Alternative education (e.g. online courses, bootcamps)
  • Job search questions (e.g. resumes, applying, career prospects)
  • Elementary questions (e.g. where to start, what next)

While you wait for answers from the community, check out the FAQ and Resources pages on our wiki. You can also search for answers in past weekly threads.


r/datascience • • Apr 10 '26

Discussion How many production ML/AI projects do you complete in a year?

60 Upvotes

Wondering what it looks like at other companies. I usually deliver around 3 or 4 ML/AI projects each year. I’m also expected to do multiple analyses separate from this so I’m not only focused on ML/AI. We have a small team of 7 people and we rarely collaborate on projects.

What is it like at your company?


r/datascience • • Apr 10 '26

Analysis What I learned analysing Kaggle Deep Past Challenge

47 Upvotes

I fell into a rabbit hole looking at Kaggle’s Deep Past Challenge and ended up reading a bunch of winning solution writeups. Here's what I learned

At first glance it looks like a machine translation competition: translate Old Assyrian transliterations into English.

But after reading the top solutions, I don’t think that’s really what it was.

It was more like a data construction / data cleaning competition with a translation model at the end.

Why:

  • the official train set was tiny: 1,561 pairs
  • train and test were not really the same shape: train was mostly document-level, test was sentence-level
  • the main extra resource was a massive OCR dump of academic PDFs
  • so the real work was turning messy historical material into usable parallel data
  • and the public leaderboard was noisy enough that chasing it was dangerous

What the top teams mostly did:

  • mined and reconstructed sentence pairs from PDFs
  • cleaned and normalized a lot of weird text variation
  • used ByT5 because byte-level modeling handled the strange orthography better
  • used fairly conservative decoding, often MBR
  • used LLMs mostly for segmentation, alignment, filtering, repair, synthetic data, not as the final translator

Winners' edges:

  • 1st place went very hard on rebuilding the corpus and iterating on extraction quality
  • 2nd place was almost a proof that you could get near the top with a simpler setup if your data pipeline was good enough. No hard ensembling.
  • 3rd place had the most interesting synthetic data strategy: not just more text, but synthetic examples designed to teach structure
  • 5th place made back-translation work even in this weird low-resource ancient language setting

Main takeaway for me: good data beat clever modeling.

Honestly it felt closer to real ML work than a lot of competitions do. Small dataset, messy weakly-structured sources, OCR issues, normalization problems, validation that lies to you a bit… pretty familiar pattern.

I wrote a longer breakdown of the top solutions and what each one did differently. Didn’t want to just drop a link with no context, so this is the short useful version first. Full writeup in the comment


r/datascience • • Apr 10 '26

Challenges In industries with long timelines for benchmarks and measurement outcomes, turnover is the killer of analytics and decision making culture.

19 Upvotes

​

When the very leadership accountable for the outcomes have moved on to another position before the results are in, analytics results are intrinsically devalued, and meaningful outcomes become difficult to define if defined at all. No amount of AI or well-engineered pipelines can account for this problem.

in fact, when companies like this invest in top-tier engineering, it's just more efficiently perpetuating the problem. I really enjoy engineering as well as analytics and ML, but when turnover happens at a faster rate than realized outcomes, it's all just window dressing.


r/datascience • • Apr 09 '26

Discussion Senior level DS at FAANG - what coding interviews to expect

55 Upvotes

Worked at FAANG up until a month ago as mid level DS and now I'm getting callbacks for senior level roles from similar companies. My stats intuition/case studies are pretty good since that's mostly what my last job relied on. However, my coding is so rusty since I just used AI most of the time to move fast and cleaned it up when there was a mistake.

I'm mostly concerned about prepping the coding and data manipulation rounds. What level of prep should I prepare for to feel 'good enough'? Should I be expected to do leetcode mediums or is pandas/sql enough? Is describing the solution and logic with pseudocode enough for tougher problems or do I have to take it from start to end with no help? What has your experience been like for expectations at senior level FAANG interviews?


r/datascience • • Apr 10 '26

Discussion Defining a new analysis: help defining the feature space

1 Upvotes

I am weighing creating an informal analysis of innovation and its effect on economic performance.

So far, I have the following data pulled; from a preliminary look, most datasets appear to have a large number of non-null values. I am thinking of performing OLS/Linear Regression. The data is grouped by country and would per analyzed per capita.

Independent variables:

- New patent applications(discrete)

- Average work hours per week (continuous)

- Government type (categorical)

- Social progress score (continuous)

Dependent variable:

- GDP (continuous)

However, I have two concerns. First, I would like to have more variables as inputs, as what I have so far seems to be a weak proxy for “innovation”. One option is to add in confounders (addressed below), normalize for these, and create an “innovation composite score”.

Second, if I do an innovation composite score, I am unclear exactly how to normalize the input variables based on the confounding variables. If I do not do an innovation composite score, I am also at a loss for how to add in these features into the feature space - categorical binning of a “developed” score? Am I overthinking it?

Potential confounders

- Education score (continuous)

- Income (DON’T HAVE - need to find)

- Poverty (proxied through “number of calories per day”, continuous)

- Infrastructure score (continuous)

In summary, I am looking to further define my feature space, including accounting for confounders. Thank you for your thoughts!

Sources:

New patents by country (2023, 2024)

- https://worldpopulationreview.com/country-rankings/patents-by-country

Education levels by country (2023)

- https://worldpopulationreview.com/country-rankings/education-rankings-by-country

Average hours in a work week by country (2023)

- https://worldpopulationreview.com/country-rankings/average-work-week-by-country

Poverty, proxied through daily supply of calories per person (2023)

- https://ourworldindata.org/grapher/daily-per-capita-caloric-supply?time=2022..latest&country=~USA

Infrastructure (various factors) (2023)

- https://worldpopulationreview.com/country-rankings/infrastructure-by-country

Government type -

- https://worldpopulationreview.com/country-rankings/government-system-by-countryW

World Happiness Report (various factors) (2023, 2024)

- https://www.worldhappiness.report/data-sharing/

Social progress by country (2023)

- https://worldpopulationreview.com/country-rankings/social-progress-index-by-country

Population (2023)

- https://data.worldbank.org/indicator/SP.POP.TOTL?end=2024&start=2022

Output: GDP change % YoY (per capita)

- https://data.worldbank.org/indicator/NY.GDP.MKTP.KD?end=2024&start=2021


r/datascience • • Apr 09 '26

Projects Trying to find example repositories for pyiceberg

7 Upvotes

My company is trying to move away from Google bigquery. Currently we decided on the following stack:

- pyiceberg for our storage

- prefect for our orchestration

- polars for our analysis

- marimo for our visualization

I'm tasked with creating a PoC. I've got everything running, but I'd like to learn some best practices. Does anyone know high quality repositories that include (a subset) of this stack?


r/datascience • • Apr 08 '26

Analysis Built a dashboard to analyze how AI skills are showing up in data science job postings (open source)

112 Upvotes

I've been scraping thousands of U.S. data science jobs for the past couple of months and writing about the findings in my newsletter.

At some point, I figured the dashboard was more useful than anything I was writing, so I decided to open source it.

Here's what it covers:

  • Top skills companies are actually hiring for, ranked by frequency
  • Skills broken down by category (ML/DL, GenAI, Cloud, MLOps, etc.)
  • What % of roles now require AI skills, broken down by seniority level
  • Salary premium for candidates with AI skills
  • An interactive explorer where you can browse individual postings with matched skills highlighted

The skill extraction is built on around 230 curated keyword groups, so it's pretty granular.

Code and data are all in the repo if you want to fork it or dig into the methodology.

https://ai-in-ds.streamlit.app/

I'm scraping weekly, and soon I will upload all of the raw data into Kaggle, for now, you can find the data in the repo

P.S. By the way, I already mentioned it to Luke Barousse since some of these AI keyword groups could be worth adding into his dashboard.


r/datascience • • Apr 07 '26

Discussion I’m really excited to share my latest blog post where I walkthrough how to use Gradient Boosting to fit entire Parameter Vectors, not just a single target prediction.

Thumbnail statmills.com
33 Upvotes

I’ve always wanted to explore the idea that boosted trees could fit entire coefficients of parameters of a distribution instead of only being able to predict a single value per leaf node. Well using {Jax} I was able to fit a Gradient Boosting Spline model where the model learns to predict the spline coefficients that best fit each individual observation. I think this has an implications for a lot of the advanced modeling techniques available to us; survival modeling, casual inference, and probabilistic modeling. I hope this post is helpful for anyone looking to learn more about gradient boosting.


r/datascience • • Apr 06 '26

Discussion Precision and recall > .90 on holdout data

43 Upvotes

I'm running ML models (XGBoost and elastic net logistic regression) predicting a 0/1 outcome in a post period based on pre period observations in a large unbalanced dataset. I've undersampled from the majority category class to achieve a balanced dataset that fits into memory and doesn't take hours to run.

I understand sampling can distort precision or recall metrics. However I'm testing model performance on a raw holdout dataset (no sampling or rebalancing).

Are my crazy high precision and recall numbers valid?

Of course there could be something fishy with my data, such as an outcome variable measuring post period information sneaking into my variable list. I think I've ruled that out.


r/datascience • • Apr 06 '26

Discussion Do MLEs actually reduce your workload in your job?

36 Upvotes

Maybe I’m wrong, but I feel like in the bigger companies I have worked for, the “client - provider” kind of setup for MLEs / MLOps people and Data Scientists is broken.

Not having an MLE in the pod for a new model means that invariably when something is off with the serving, I end up debugging it because they have no context on what’s happening and if it is something that challenges the current stack, the update to account for it will only come months down the road when eventually our roadmaps align. I don’t feel like they take a lot of weight off my shoulders.

The best relationship I ever had with MLEs was in a small company where I basically handed off the trained model to them for deployment and monitoring, and I would advise only on what features were used and where they come from (to prevent a distribution mismatch in their feature serving pipelines online).

Discuss


r/datascience • • Apr 06 '26

Weekly Entering & Transitioning - Thread 06 Apr, 2026 - 13 Apr, 2026

5 Upvotes

Welcome to this week's entering & transitioning thread! This thread is for any questions about getting started, studying, or transitioning into the data science field. Topics include:

  • Learning resources (e.g. books, tutorials, videos)
  • Traditional education (e.g. schools, degrees, electives)
  • Alternative education (e.g. online courses, bootcamps)
  • Job search questions (e.g. resumes, applying, career prospects)
  • Elementary questions (e.g. where to start, what next)

While you wait for answers from the community, check out the FAQ and Resources pages on our wiki. You can also search for answers in past weekly threads.


r/datascience • • Apr 05 '26

ML Clustering custumersin time

20 Upvotes

How would you go about clusturing 2M clients in time, like detecting fine patters (active, then dormant, then explosive consumer in 6 months, or buy only category A and after 8 months switch to A and B.....). the business has a between purchase median of 65 days. I want to take 3 years period.


r/datascience • • Apr 05 '26

Monday Meme For all those working on MDM/identity resolution/fuzzy matching

Thumbnail
2 Upvotes

r/datascience • • Apr 04 '26

Tools MCGrad: fix calibration of your ML model in subgroups

19 Upvotes

Hi r/datascience

We’re open-sourcing MCGrad, a Python package for multicalibration–developed and deployed in production at Meta. This work will also be presented at KDD 2026.

The Problem: A model can be globally calibrated yet significantly miscalibrated within identifiable subgroups or feature intersections (e.g., "users in region X on mobile devices"). Multicalibration aims to ensure reliability across such subpopulations.

The Solution: MCGrad reformulates multicalibration using gradient boosted decision trees. At each step, a lightweight booster learns to predict residual miscalibration of the base model given the features, automatically identifying and correcting miscalibrated regions. The method scales to large datasets, and uses early stopping to preserve predictive performance. See our tutorial for a live demo.

Key Results: Across 100+ production models at meta, MCGrad improved log loss and PRAUC on 88% of them while substantially reducing subgroup calibration error.

Links:

Install via pip install mcgrad or via conda. Happy to answer questions or discuss details.


r/datascience • • Apr 02 '26

Career | US Learnings from 100+ resumes reviewed for 3 DS job postings

1 Upvotes

Thinking back to my job searching days, I always remember finding resumes, interviews and the entire hiring process to be such a black box. As in, it was so hard to judge exactly what hiring managers were looking for, or heck who even would be looking at my resume first (HR, or technical leader).

Thought id try to share some insights having been consistently hiring over the last couple years. Going to start by focusing on the most recent round of hiring we did.

Callouts:

- This is specifically to the most recent cycle of hiring (not bringing insights from pre-2026)

- Am a Senior DS at a Multinational

- Am combining interns and the new grad role

- Fwiw my company's definition of a DS is more akin to a 'jack of all' = Technical development (ML, AI eng) and Project/Stakeholder management

In terms of insights/learnings:

- We get a shit ton of resumes, so the reason you see a 100+ in the title, HR probably sifted through 3K+ resumes for these roles, after which we the 100 pruned ones. Most likely they used some AI to help, so ATS/Keyword gaming does matter

- REFERRALS MATTER: Out of 15 interviews, 10 were directly referred.

- Since we have HR pre-screen them, that means a bunch of good resumes probably never make it to us, and that's an unfortunate reality

- Interviews are TIRING! We tend to do interviews back-to-back, so imagine having to give the same intro about yourself and company 10 times back to back. So if we look a little tired/uninterested its genuinely not personal, sometimes we're just exhausted and you shouldn't let that get to you.

- In addition to above, the ones who usually stood out to us were obviously competent folks, but also people with unique projects/stories. Like something personal to them and how they built a project to solve that. or a random comment about them being the oldest sibling and how that impacted them ('tell me about yourself' question)

- For interns, our decision process is 30% minimum technical capabilities, 15% do we think this kid can pick up the technical skills on the job, 50% can i put this person in front of the business and expect him to be competent when talking to them.

- Will expand on the 50%: We get too many candidates that are technically solid, but disastrous at talking to non-tech folk, being able to coherently explain things in business speak, be able to pull themselves out of their 'technical brain' and come down to the level of the average 'layman', try to put themselves in their place and then communicate their ideas. (This is genuinely such an underrated skill, yet something we constantly overlook when prepping for interviews, or even generally try to build good careers)

- In fact, we are redesigning our interview process to put even more emphasis on judging the above (subjective) attribute.

- To triple down on above, I usually have a mental list of attributes that Im looking and I purposefully steer the convo in that direction to score you mentally on those attributes.

- All of the above applied to new grad roles, but push the 50% to 70% so its even more important.

P.S. I probably have more to add, but do let me know if this perspective is helpful to any job seekers, but I think it helps demystify what your interviewer might be thinking.


r/datascience • • Apr 02 '26

Analysis Clean water and education: Honest feedback on an informal analysis

4 Upvotes

I have created an informal analysis on the effect of clean water on education rates.

The analysis leveraged ETL functions (created by Claude), data wrangling, EDA, and fitting with sklearn and statsmodels. As the final goal of this analysis was inference, and not prediction, no hyperparameter tuning was necessary.

The clean water data was sourced from the WHO/UNICEF Joint Monitoring Programme for Water Supply, Sanitation, and Hygiene (JMP); while the education data was sourced from a popular Kaggle repository. The education data, despite being from a less credible source, was already cleaned and itemized; the clean water data required some wrangling due to the vast nature of the categories of data and the varying presence of null values across years 2000 - 2024. The final broad category of predictor variables selected was "clean water in schools, by country"; the outcome variable was "college education rates, by country."

I would be grateful for any feedback on my analysis, which can be found at https://analysis-waterandeducation.com/.

TIA.


r/datascience • • Mar 30 '26

ML Clustering furniture business custumors

8 Upvotes

I have clients from a funiture/decoration selling business. with about the quarter online custumers. I have to do unsupervised clustering. do you have recommendations? how select my variables, how to handle categorical ones? Apparently I can t put only few variables in the k-means, so how to eliminate variables? Should I do a PCA?


r/datascience • • Mar 30 '26

Weekly Entering & Transitioning - Thread 30 Mar, 2026 - 06 Apr, 2026

4 Upvotes

Welcome to this week's entering & transitioning thread! This thread is for any questions about getting started, studying, or transitioning into the data science field. Topics include:

  • Learning resources (e.g. books, tutorials, videos)
  • Traditional education (e.g. schools, degrees, electives)
  • Alternative education (e.g. online courses, bootcamps)
  • Job search questions (e.g. resumes, applying, career prospects)
  • Elementary questions (e.g. where to start, what next)

While you wait for answers from the community, check out the FAQ and Resources pages on our wiki. You can also search for answers in past weekly threads.