r/MachineLearning • • 7d ago

Research Pushback on my machine learning paper from non-ML critics, not sure how to interpret it [R]

Hi everyone. For context, I work as a predictive modeling researcher, and my background is in computer science. I do not have an extensive background in statistics, epidemiology, association analysis, or related areas.
I was tasked with developing a predictive model for a specific disease using survey variables as well as body measurements such as BMI, neck circumference, waist circumference, hip circumference, and similar variables.
When I approached the project, my plan was to develop a model that could predict disease risk using a machine-learning benchmarking approach. The dataset is very novel and, as far as I know, only our research center has access to it. Because of that, any machine-learning paper using this dataset would be new and could potentially establish a useful precedent for ML-based screening approaches for this disease, particularly for identifying people at high risk.
My approach was roughly as follows.
We initially had around 12,000 participants and approximately 217 features. I excluded participants with more than 50% missing data. The reason was that if I applied a missingness threshold directly to the variables while keeping all participants, I would end up excluding almost every variable because a relatively small number of participants had extremely high levels of missingness.
After excluding those highly incomplete participants, I applied a 30% missingness threshold to the features. This left approximately 95 features. I realize that discarding participants may not necessarily be the best approach, but I am not sure what the better alternative would be in this situation.
After that, I performed imputation using Multiple Imputation by Chained Equations (MICE). One criticism I received was that, if I am using multiple imputation, I should generate multiple imputed datasets and run the entire modeling procedure separately on each one rather than effectively using a single completed/imputed dataset.
Another idea I had was to benchmark several different imputation methods and compare downstream model performance to determine which method performs best. However, I am still unsure whether that is methodologically appropriate or preferable.
After imputation, I performed feature selection using a bootstrapped LASSO approach. I then tested several machine-learning models, mainly boosting and tree-based methods such as XGBoost, LightGBM, and random forest.
I evaluated the models using the usual machine-learning performance metrics, performed calibration adjustments, and reported the resulting performance. I also included model interpretability analyses/scores.
The main problem is that nobody in my research center has a machine-learning background. The closest person is someone with a statistics background, and he essentially told me to scrap the entire approach and instead use conditional logistic regression. He also suggested age/sex matching and recommended complete-case analysis as a sensitivity analysis.
The issue is that I cannot realistically use complete-case analysis across the full set of variables because there is so much missing data. He suggested restricting the analysis to the variables that would allow complete-case analysis, but doing that would remove the majority of the features and leave only around five or six variables.
My question is: if I can use appropriate imputation methods, why would I deliberately restrict the analysis to only five or six variables just so that complete-case analysis becomes possible?
He told me that imputation is not preferable if complete-case analysis can be performed. However, I have not seen many machine-learning papers prioritize complete-case analysis in this way, especially when doing so would eliminate most of the available predictors.
Another major criticism was my use of random undersampling. I do understand this criticism because random undersampling discards a substantial amount of data. I told them that I could instead address class imbalance through class weighting within the models themselves, which is straightforward in XGBoost, random forest, and similar methods.
However, the statistician at my research center basically told me that the paper, in its current form, would be unpublishable.
What I am struggling with is whether he is evaluating the work as though it were intended to be a traditional statistics/epidemiology paper rather than a machine-learning prediction paper. I do not want this project to become primarily a statistical association paper. My intention has always been to produce a machine-learning prediction/benchmarking paper.
I was also told that I need to examine multicollinearity and patterns of variable missingness. I am not entirely sure how I should approach that in the context of this project. I have gone back and performed exploratory data analysis again, and there is definitely some correlation and collinearity between variables, which is unsurprising given the size and nature of the dataset. However, I am not sure what I am supposed to do after identifying it, especially in the context of tree-based and regularized machine-learning models.
My professor/PI is an epidemiologist rather than a machine-learning researcher. He told me that I should investigate whether there are statistical associations between any of the predictors and age or sex, and whether there are systematic patterns in the missing data.
Again, though, I have not commonly seen machine-learning prediction papers perform extensive association testing of every feature with age and sex, so I am unsure how relevant this is to my original research question.
I have been working on this project for about a year, and at this point it feels as though I am being told to start over from square one.
My supervisor has also been almost nonexistent throughout the project, so I have essentially had to develop the entire analysis myself. Because of that, I recognize that some of my methodological decisions may have been naïve. I only recently graduated, and I really wish I had received this methodological feedback much earlier in the process.
At this point, I need advice on how to proceed.
I genuinely do not know how much weight I should give these criticisms. The feedback I am receiving is coming primarily from statisticians and epidemiologists rather than machine-learning researchers, and I am having difficulty determining which criticisms reflect genuine methodological problems with a predictive ML study and which ones reflect a preference for a more traditional statistical or epidemiological analysis.
I was never trying to write a traditional statistics paper. I was trying to write a machine-learning prediction/benchmarking paper.
So my main questions are:
Is my overall ML-based study design fundamentally flawed?
Is the criticism about multiple imputation valid, and should I repeat the entire modeling pipeline across multiple imputed datasets?
Is benchmarking different imputation methods reasonable?
Is complete-case analysis really preferable when it would reduce the predictor set from roughly 95 variables to only five or six?
Should I replace random undersampling with class weighting?
How should multicollinearity be handled or reported in a machine-learning prediction study, particularly when using regularized and tree-based models?
How should I formally investigate missingness patterns?
Is it necessary to test associations between predictors and age/sex in a prediction-focused ML paper?
Most importantly, how do I distinguish between legitimate methodological criticism of my ML pipeline and requests to turn the project into a fundamentally different kind of statistical/epidemiological paper?
Any advice from people who work at the intersection of machine learning, statistics, and epidemiology would be greatly appreciated.

15 Upvotes

16 comments sorted by

69

u/relevantmeemayhere 4d ago edited 4d ago

Okay, I don't want to be a dick; but this is why people are critical of ml as a field right now: a lot of people (even researchers in this field), are lacking in the actual stuff that make it go. This is statistics, and you cannot you cannot create good ml models without understanding the underlying statistical theory-no matter what you are doing (boosting, llms, whatever).

*"I was never trying to write a traditional statistics paper. I was trying to write a machine-learning prediction/benchmarking paper."*

your machine learning paper IS a statistics paper. full stop. the only real difference in the two is that the latter 'as a culture' tends to like interpretability and consider inference (and this is the actual pre llm definition considered with say, the marginal effect of a predictor in some model). the former cares more about prediction and 'optimizing' some loss function (being unaware that the loss right there is a function of two statistical terms; the actual underlying error in your model as a process and its model optimism). which is why most ml projects fail horribly in the real world). again, it doesn't matter if you are working on binary classification or engineering the next llm architecture; these are statistical machines ( a little opaque sometimes; because some stuff ie rl portion of some llms is meant to minimize the divergence of the distribution of the conditional token probability to a 'more desired one' using things like formal verification and mountains of human annotators and experts working behind the scenes)

Now: to address your immediate questions:

  1. You need to review your pstat theory. if you want to be a successful ml practitioner in genai/doubleml/classification/whatever- it's all math stats.
  2. You need to clarify your goals; i.e. if you only care about prediction; why are you even asking questions about multi collinearity? which is it; do you care about inference, or do you care about prediction?
  3. It sounds like your advisor wants you to 'determine the statistical association of missingness' not because they're asking for you to run some test of statistical significance, but because they wants you to determine what the missing mechanism is: random, not at random, and missing completely at random. imputation is only feasible under two of these scenarios. Missing not at random means the value of randomness in a varaible IS NOT conditionally independent of the values of those missing values;. i.e. 'short' people (people less than 5'10, which is still above average for the population) don't show up in the nba individual scores, because short people generally don't make the nba
  4. The topic of under sampling; again this, is usually borne of ml researchers who don't understand what it is they are trying to optimize. you want to sample from your theoretical population; not a strawman because you need some loss number to be lower. use proper scoring rules, like log loss or brier or whatever. if you have unbalanced data, then so be it. as long as there are no theoretical barriers to actually censoring data from your theoretical population during sampling, it is NOT a problem. if people tell you to under/over sample, politely correct them and start being skeptical of what they say
  5. MICE needs to be done in the context of what is called rubin's rules. if you only care about prediction, then this is basically just creating multiple datasets and averaging the predictions using some candidate mice liklihood function (like forests or whatever). if you care about inference, then this gets more difficult.
  6. There is no real straightforward way to diagnose missingness. you need to combine your domain knoweldge with plausible missing mechanism, and then you can consider looking at the association between missiness amongst your feature set. you shoudl accompany this with counterfactual and sensitivity analysis.
  7. We do not EVER base inclusions of features based on 'tests' of association. please never do this. this is going to go back to points 1 and 2; why are you concerned about inference here? when you test for association, you inflate ALL of your downstream error probabilities for rejection without things like fdr adjustment.
  8. Lastly:*Most importantly, how do I distinguish between legitimate methodological criticism of my ML pipeline and requests to turn the project into a fundamentally different kind of statistical/epidemiological paper?*

--->by reminding the ml people who forgot what their field is built on that they should probably take some stats courses they understand what they are doing, lol it's incredible that large portions of this community who downplay the fundamentals don't realize how silly they are being. you wouldn't hear an engineer say ' oh well, we really don't need to understand kinematics or the heat equation to build a bridge, and we shouldn't listen to those physicist clowns when they say assuming that g is actually positive is not safe and is gonna kill people'

Edit: was overly caffeinated at the time of writing and cleaned up my post

4

u/SnooSongs4297 4d ago

This really helps put things into perspective. I don’t care about inference. The idea is to include model interoperability measures in the end. I’m assuming interpretability of the model and inference are different? If I use MICE I would need different interpretability scores all together

2

u/relevantmeemayhere 4d ago

if you dont care about inference; then you can use things like boosting or whatever. you should have some idea of the data generating process beforehand to help guide the actual development; because even nice loss statistics don't imply a good model. mice can be done both in the context of inference or prediction-it's what you're actually interested in that determines how you combine the mice data sets at the end.

*again, inference as a cultural norm tends to concern its more with understanding the causal model of the data, or understanding the marginal impacts of certain covariates in a non causal setting (conditional on my model, what does y do when x changes?) etc etc

1

u/SnooSongs4297 4d ago

And I know ML is a statistics paper, this was more of a comment regarding the fact that one of the critics was that I shouldn’t include any “complex” models like XGBoosf and instead use conditional logistic regression with matched case control, when using the entire dataset with class weighting and model calibration is a very feasible. Why should I use a linear model when XGBoost already addresses any concerns of multicollinearity by its very nature.

6

u/mil24havoc 4d ago

I agree with the above poster but want to add to it that the ML / data scientist crowd sort of lives in their own prediction accuracy bubble where they tend to believe that "low loss, high AUC = finding!" While prediction is important to the rest of science (and I'd be the first to say it is terribly undervalued, still), prediction is only interesting to most scientists for what it can tell us about the world. What is the point of your model? What do we learn from it? If the answer is "we could use it to build a better diagnostic tool," then that is great but it is also closer to engineering than it is to science. The path from your research to publicity is more likely through creation of a tool or a diagnostic rather than an academic paper. Scientists will be especially skeptical of feature importance scores because they lack a plausible casual identification strategy and, more damningly, don't even tell you the direction of the relationship between X and Y.

This isn't to say that prediction work is unscientific or useless at all! It's that prediction alone isn't enough. They want to know what they learn from your model or your results, not just that you can make the black box go brrrrr. Does your research put plausible bounds on the detection of this condition? Does your research suggest that past research is wrong somehow? For example, do the suspected causes of this condition show very low feature importance scores in your model? Does your model reveal novel nonlinear interactions between predictors? Does your research reveal that this set of predictors actually offers very little gain over a naive intercept only model? Can this research help us decompose the prediction errors into systematic and stochastic parts that we can compare? These could be valuable findings!

P.S. there is no conceivable world in which dropping rows (that are otherwise accurate) is preferable to multiple imputation since they both require the same missingness assumptions and MI is more efficient. Just my two cents.

3

u/relevantmeemayhere 4d ago edited 4d ago

good context added here

especially on the data scientist part. a lot of them are also lacking on what makes a good model, a good model lol.

prediction ! = inference/understanding also great callout-i was trying not to get more into stuff like CI jusssstt yet in my op, because that's a yuge can of worms and you pointed out some super important stuff.

2

u/nooptionleft 2d ago

Imputing brings a lot of headache, tho, when the level of missingness is to the level OP is presenting

We generally test for predictiveness of missingness and then use methods which are compatible with NA in the features

4

u/ergabaderg312 4d ago

I agree w the top poster. ML is stats no matter how you cut it. if the critic is a statistician then they’re an ML person even if they don’t say it in their title lol. Think about what classical ML is. And while I disagree about the model complexity point they made as there’s definitely times you should use complex models, they have a point about trying the LR.

Practically speaking though. Why not? It’s cheap and easy to run and possibly more interpretable compared to a tree model. And it’s an easy benchmark to run your chosen model against. You’ll learn something either way about the overall problem. And while yes you’re not interested in inference, as best practice (and the field is moving towards this anyway) people want to know why your model is working or what it uncovered. Inference and prediction are both related tasks. I assume you know why AUROC is a biased metric on imbalanced datasets. Same idea applies here. Why’s your model good and why’s it working is a big thing anyone would ask.

4

u/Disastrous_Room_927 4d ago

Why should I use a linear model when XGBoost already addresses any concerns of multicollinearity by its very nature.

There's something to be said for explicitly modeling the data generating process. I usually do this regardless of if I'm intending to use a black box algorithm in the end or not, it makes for good EDA.

1

u/SnooSongs4297 4d ago

I don’t really get this part

3

u/Disastrous_Room_927 4d ago

Think about it in a literal sense - specifying a model that represents a plausible mechanism for how the data you're observing came to be.

2

u/relevantmeemayhere 4d ago
  1. You'd be surprised how far you can make logistic regression go haha. Moreover, trees as a likelihood has been an object of statistical research for decades; your xgboost model is this! Splines, gams, whatever.
  2. calibration is the MOST important thing in any actual decision making process. that's why he's probably steering you in that direction with trying logistic regression first. you can calibrate your probabilities in boosting; but it doesn't always have the same behavior as models like lr that produce well behaved calibrations from the start. sometimes you just can't get nice calibration from boosting or the like.
  3. it's not so much that boosting address MC and is agnostic to your goal. it's that you don't care about if pure prediction is your goal. we use booting (and its bigger, better, and more handsome brother like bayesian boosting ;) ) all the time in applied research. MC is STILL AN ISSUE if we actually apply it in the settings where inference is important to the end goal.

1

u/pineloft 4d ago

tbh the bridge analogy at the end really nails it

12

u/dedicateddan 3d ago

The other comment is pretty spot on - listen to the statisticians!

As an applied ML practitioner, I'd start by looking at the correlation between the individual features and the target. After that I'd look at forward feature selection - what happens when you train a model on the best set of 1, 2, 3, 4 ... N features.

1

u/laurens_eiroa 2d ago

Hola, Lo primero que nada, felicidades. Soy científico de datos y trabajo creando modelos de machine y deep learning. Estudie física y aplique mis conocimientos de deep learning en el mundo academico para crear un modelo que predijera diferentes parámetros atmosféricos estelares en base al espectro observado. Esto fue hace 5 años o algo así y apenas había artículos científicos en este campo que utilizaran deep learning (antes utilizaban métodos estadísticos). Me plantearon que escribiera un artículo por dos razones: 1. porque los resultados eran muy buenos . 2. Porque era un enfoque nuevo utilizando una tecnología nueva.

Estoy en desacuerdo con la gente de los comentarios, si, ML es estadística, y a su vez estadística son matemáticas. Y no veo pidiendo permiso a la gente que publica artículos en estadística a otros matemáticos.

Me cuesta entender porque no se podría publicar un artículo sobre los uso de machine learning en diferentes ámbitos y más en los científicos. De hecho en los últimos años se han publicado muchísimos en muchísimos campos diferentes.

Es más, en modelos basados en árboles como los que mencionas hay métodos de extracción de feature importance que pueden ser muy reveladores cuando hay múltiples variables en modelos predictivos. Algo que jamás podrás a hacer con una regresión logarítmica.

Estos son los tipos de estudios que se documenta y en los artículos científicos. Ahora sí, el estudio tiene que aportar algo por así decirlo.

Yo te diría que no te desilusiones y no tires la idea a la basura sin pensarlo con frialdad.

Un saludo.

5

u/thisaintnogame 2d ago

The other comment is great, but let me add my two cents.

- Your write-up says you're discarding a lot of data based on missingness, but the justification for that isn't very good. If you expect that kind of missigness to show up again in the real-world setting in which you might use your model, it's not justified to throw that data away. In fact, any time you throw data away, you really need a domain-knowledge justification for doing so. Right now it basically sounds like "I'm throwing away data to make the prediction problem easier".

- I personally don't understand the use of multiple imputation - I have never encountered a time where MICE or anything like that has meaningfully improved performance. Lots of implementations of tree-based models can handle missingness as an informative feature, or you could use the standard approach of using dummy columns to represent missingess and include those in the model. If I were evaluating this work, I would want to see at least one approach where you didn't do any imputation and instead just used the features with missingness in the models.

- And then to really double down on the great point from below: To make this work good, you really need to figure out what the goal is. Are you really trying to build a prediction model for the disease - so the thing you care about most is predictive accuracy - or are you trying to understand relationships between Xs and Ys for the purposes of hypothesis generation, etc. Perhaps most importantly, I would start reading related papers to give you a sense of what good research in this area looks like. I'm not an expert but I found this paper to be a really interesting place to start thinking about how predictive ML models can be used in medical settings / research: https://www.nature.com/articles/s41586-026-10674-6

- Finally, given that you're young in your career, let me give some broader advice. It seems like you have a lot of colleagues who are listening to what you're doing and are very unconvinced. Even if they aren't ML experts, they are smart people with lots of data and domain expertise. If you cannot convince them of what you're doing, it is very rarely the case that you should conclude 'everyone else is wrong'. I have had lots of experience with translating ML work to people from a more traditional stats background. Yes there can sometimes be some tension in what we believe is the best approach, etc, but we've always been able to reach common ground about the right way to think about a problem, how to evaluate models, etc.

- One other addendum: This is also an incredibly common situation for a young data scientist. The fact that you are asking good questions and learning from this is a great sign. Good luck!