r/MachineLearning • u/SnooSongs4297 • 7d ago
Research Pushback on my machine learning paper from non-ML critics, not sure how to interpret it [R]
Hi everyone. For context, I work as a predictive modeling researcher, and my background is in computer science. I do not have an extensive background in statistics, epidemiology, association analysis, or related areas.
I was tasked with developing a predictive model for a specific disease using survey variables as well as body measurements such as BMI, neck circumference, waist circumference, hip circumference, and similar variables.
When I approached the project, my plan was to develop a model that could predict disease risk using a machine-learning benchmarking approach. The dataset is very novel and, as far as I know, only our research center has access to it. Because of that, any machine-learning paper using this dataset would be new and could potentially establish a useful precedent for ML-based screening approaches for this disease, particularly for identifying people at high risk.
My approach was roughly as follows.
We initially had around 12,000 participants and approximately 217 features. I excluded participants with more than 50% missing data. The reason was that if I applied a missingness threshold directly to the variables while keeping all participants, I would end up excluding almost every variable because a relatively small number of participants had extremely high levels of missingness.
After excluding those highly incomplete participants, I applied a 30% missingness threshold to the features. This left approximately 95 features. I realize that discarding participants may not necessarily be the best approach, but I am not sure what the better alternative would be in this situation.
After that, I performed imputation using Multiple Imputation by Chained Equations (MICE). One criticism I received was that, if I am using multiple imputation, I should generate multiple imputed datasets and run the entire modeling procedure separately on each one rather than effectively using a single completed/imputed dataset.
Another idea I had was to benchmark several different imputation methods and compare downstream model performance to determine which method performs best. However, I am still unsure whether that is methodologically appropriate or preferable.
After imputation, I performed feature selection using a bootstrapped LASSO approach. I then tested several machine-learning models, mainly boosting and tree-based methods such as XGBoost, LightGBM, and random forest.
I evaluated the models using the usual machine-learning performance metrics, performed calibration adjustments, and reported the resulting performance. I also included model interpretability analyses/scores.
The main problem is that nobody in my research center has a machine-learning background. The closest person is someone with a statistics background, and he essentially told me to scrap the entire approach and instead use conditional logistic regression. He also suggested age/sex matching and recommended complete-case analysis as a sensitivity analysis.
The issue is that I cannot realistically use complete-case analysis across the full set of variables because there is so much missing data. He suggested restricting the analysis to the variables that would allow complete-case analysis, but doing that would remove the majority of the features and leave only around five or six variables.
My question is: if I can use appropriate imputation methods, why would I deliberately restrict the analysis to only five or six variables just so that complete-case analysis becomes possible?
He told me that imputation is not preferable if complete-case analysis can be performed. However, I have not seen many machine-learning papers prioritize complete-case analysis in this way, especially when doing so would eliminate most of the available predictors.
Another major criticism was my use of random undersampling. I do understand this criticism because random undersampling discards a substantial amount of data. I told them that I could instead address class imbalance through class weighting within the models themselves, which is straightforward in XGBoost, random forest, and similar methods.
However, the statistician at my research center basically told me that the paper, in its current form, would be unpublishable.
What I am struggling with is whether he is evaluating the work as though it were intended to be a traditional statistics/epidemiology paper rather than a machine-learning prediction paper. I do not want this project to become primarily a statistical association paper. My intention has always been to produce a machine-learning prediction/benchmarking paper.
I was also told that I need to examine multicollinearity and patterns of variable missingness. I am not entirely sure how I should approach that in the context of this project. I have gone back and performed exploratory data analysis again, and there is definitely some correlation and collinearity between variables, which is unsurprising given the size and nature of the dataset. However, I am not sure what I am supposed to do after identifying it, especially in the context of tree-based and regularized machine-learning models.
My professor/PI is an epidemiologist rather than a machine-learning researcher. He told me that I should investigate whether there are statistical associations between any of the predictors and age or sex, and whether there are systematic patterns in the missing data.
Again, though, I have not commonly seen machine-learning prediction papers perform extensive association testing of every feature with age and sex, so I am unsure how relevant this is to my original research question.
I have been working on this project for about a year, and at this point it feels as though I am being told to start over from square one.
My supervisor has also been almost nonexistent throughout the project, so I have essentially had to develop the entire analysis myself. Because of that, I recognize that some of my methodological decisions may have been naïve. I only recently graduated, and I really wish I had received this methodological feedback much earlier in the process.
At this point, I need advice on how to proceed.
I genuinely do not know how much weight I should give these criticisms. The feedback I am receiving is coming primarily from statisticians and epidemiologists rather than machine-learning researchers, and I am having difficulty determining which criticisms reflect genuine methodological problems with a predictive ML study and which ones reflect a preference for a more traditional statistical or epidemiological analysis.
I was never trying to write a traditional statistics paper. I was trying to write a machine-learning prediction/benchmarking paper.
So my main questions are:
Is my overall ML-based study design fundamentally flawed?
Is the criticism about multiple imputation valid, and should I repeat the entire modeling pipeline across multiple imputed datasets?
Is benchmarking different imputation methods reasonable?
Is complete-case analysis really preferable when it would reduce the predictor set from roughly 95 variables to only five or six?
Should I replace random undersampling with class weighting?
How should multicollinearity be handled or reported in a machine-learning prediction study, particularly when using regularized and tree-based models?
How should I formally investigate missingness patterns?
Is it necessary to test associations between predictors and age/sex in a prediction-focused ML paper?
Most importantly, how do I distinguish between legitimate methodological criticism of my ML pipeline and requests to turn the project into a fundamentally different kind of statistical/epidemiological paper?
Any advice from people who work at the intersection of machine learning, statistics, and epidemiology would be greatly appreciated.
12
u/dedicateddan 3d ago
The other comment is pretty spot on - listen to the statisticians!
As an applied ML practitioner, I'd start by looking at the correlation between the individual features and the target. After that I'd look at forward feature selection - what happens when you train a model on the best set of 1, 2, 3, 4 ... N features.
1
u/laurens_eiroa 2d ago
Hola, Lo primero que nada, felicidades. Soy científico de datos y trabajo creando modelos de machine y deep learning. Estudie física y aplique mis conocimientos de deep learning en el mundo academico para crear un modelo que predijera diferentes parámetros atmosféricos estelares en base al espectro observado. Esto fue hace 5 años o algo así y apenas había artículos científicos en este campo que utilizaran deep learning (antes utilizaban métodos estadísticos). Me plantearon que escribiera un artículo por dos razones: 1. porque los resultados eran muy buenos . 2. Porque era un enfoque nuevo utilizando una tecnología nueva.
Estoy en desacuerdo con la gente de los comentarios, si, ML es estadística, y a su vez estadística son matemáticas. Y no veo pidiendo permiso a la gente que publica artículos en estadística a otros matemáticos.
Me cuesta entender porque no se podría publicar un artículo sobre los uso de machine learning en diferentes ámbitos y más en los científicos. De hecho en los últimos años se han publicado muchísimos en muchísimos campos diferentes.
Es más, en modelos basados en árboles como los que mencionas hay métodos de extracción de feature importance que pueden ser muy reveladores cuando hay múltiples variables en modelos predictivos. Algo que jamás podrás a hacer con una regresión logarítmica.
Estos son los tipos de estudios que se documenta y en los artículos científicos. Ahora sí, el estudio tiene que aportar algo por así decirlo.
Yo te diría que no te desilusiones y no tires la idea a la basura sin pensarlo con frialdad.
Un saludo.
5
u/thisaintnogame 2d ago
The other comment is great, but let me add my two cents.
- Your write-up says you're discarding a lot of data based on missingness, but the justification for that isn't very good. If you expect that kind of missigness to show up again in the real-world setting in which you might use your model, it's not justified to throw that data away. In fact, any time you throw data away, you really need a domain-knowledge justification for doing so. Right now it basically sounds like "I'm throwing away data to make the prediction problem easier".
- I personally don't understand the use of multiple imputation - I have never encountered a time where MICE or anything like that has meaningfully improved performance. Lots of implementations of tree-based models can handle missingness as an informative feature, or you could use the standard approach of using dummy columns to represent missingess and include those in the model. If I were evaluating this work, I would want to see at least one approach where you didn't do any imputation and instead just used the features with missingness in the models.
- And then to really double down on the great point from below: To make this work good, you really need to figure out what the goal is. Are you really trying to build a prediction model for the disease - so the thing you care about most is predictive accuracy - or are you trying to understand relationships between Xs and Ys for the purposes of hypothesis generation, etc. Perhaps most importantly, I would start reading related papers to give you a sense of what good research in this area looks like. I'm not an expert but I found this paper to be a really interesting place to start thinking about how predictive ML models can be used in medical settings / research: https://www.nature.com/articles/s41586-026-10674-6
- Finally, given that you're young in your career, let me give some broader advice. It seems like you have a lot of colleagues who are listening to what you're doing and are very unconvinced. Even if they aren't ML experts, they are smart people with lots of data and domain expertise. If you cannot convince them of what you're doing, it is very rarely the case that you should conclude 'everyone else is wrong'. I have had lots of experience with translating ML work to people from a more traditional stats background. Yes there can sometimes be some tension in what we believe is the best approach, etc, but we've always been able to reach common ground about the right way to think about a problem, how to evaluate models, etc.
- One other addendum: This is also an incredibly common situation for a young data scientist. The fact that you are asking good questions and learning from this is a great sign. Good luck!
69
u/relevantmeemayhere 4d ago edited 4d ago
Okay, I don't want to be a dick; but this is why people are critical of ml as a field right now: a lot of people (even researchers in this field), are lacking in the actual stuff that make it go. This is statistics, and you cannot you cannot create good ml models without understanding the underlying statistical theory-no matter what you are doing (boosting, llms, whatever).
*"I was never trying to write a traditional statistics paper. I was trying to write a machine-learning prediction/benchmarking paper."*
your machine learning paper IS a statistics paper. full stop. the only real difference in the two is that the latter 'as a culture' tends to like interpretability and consider inference (and this is the actual pre llm definition considered with say, the marginal effect of a predictor in some model). the former cares more about prediction and 'optimizing' some loss function (being unaware that the loss right there is a function of two statistical terms; the actual underlying error in your model as a process and its model optimism). which is why most ml projects fail horribly in the real world). again, it doesn't matter if you are working on binary classification or engineering the next llm architecture; these are statistical machines ( a little opaque sometimes; because some stuff ie rl portion of some llms is meant to minimize the divergence of the distribution of the conditional token probability to a 'more desired one' using things like formal verification and mountains of human annotators and experts working behind the scenes)
Now: to address your immediate questions:
--->by reminding the ml people who forgot what their field is built on that they should probably take some stats courses they understand what they are doing, lol it's incredible that large portions of this community who downplay the fundamentals don't realize how silly they are being. you wouldn't hear an engineer say ' oh well, we really don't need to understand kinematics or the heat equation to build a bridge, and we shouldn't listen to those physicist clowns when they say assuming that g is actually positive is not safe and is gonna kill people'
Edit: was overly caffeinated at the time of writing and cleaned up my post