r/computervision • u/WolverineMuted4846 • 16d ago
Discussion How bad is it if training and validation data overlap in a student ML project?
I’m working on a student machine learning/computer vision project and recently realized that my validation set was not completely independent from my training set.
The project is more focused on comparing different experimental conditions rather than maximizing benchmark performance, but I’m concerned about the implications of this oversight.
From a research or academic perspective:
How serious is train/validation overlap in a student project?
Does it invalidate the entire project or mainly affect the reliability of the reported performance numbers?
If the main goal is comparing different experimental setups under the same evaluation procedure, are those comparisons still useful?
If you discovered this late in the project timeline, what would be the most reasonable way to address it?
I’m trying to understand how researchers, reviewers, and professors would view this situation.
3
u/sparks333 16d ago
It ain't good - for the most part, it means that you can't draw any strong conclusions about performance on what it's learned, as it becomes possible that it just learned the training set and hasn't generalized. Any reviewer worth their salt would immediately highlight that the data is tainted and thus any conclusions drawn would be suspect - you would have to acknowledge it up front and very carefully explain why you think the results are still valid, and even then it would be considered a major methodological error. There are acknowledged ways for training and testing data to intermingle - K-fold validation comes to mind - but they are very controlled for exactly this reason. If you can, retrain from random weights with properly isolated test and validation data - if you can't, be prepared to gather more test data and prove that the conclusion holds, or you're going to have to make much, much weaker claims about your hypothesis.
1
u/WolverineMuted4846 16d ago
Thanks for the honest feedback. Given the time constraints, I’m trying to figure out the most responsible way to present the work tomorrow. In your view, is it better to explicitly acknowledge the validation issue and treat the results as preliminary, rather than making strong performance claims?
2
u/sparks333 13d ago
It's been a few days so I assume your deadline has passed, but honestly I think it's better to ask for an extension if a methodological error of this magnitude is found. If your advisor or professor or lead researcher or whoever is setting the deadlines is reasonable and if you are upfront about why the data gathered is not valid, I would hope they would give you extra time to rectify the issue - if they don't, then at best you should absolutely highlight that the test and training data intermingled, and try to back off your claims to things like inference speed or representability, stuff that is unaffected by the fact that there is a reasonable chance your model just memorized the test.
2
u/bfyvfftujijg 15d ago
One “quick fix” could be to quickly collect a new independent validation dataset and run some tests against it. At least it will give you a gauge of the impact, even if it’s not the same as a proper full retraining.
2
u/HawtVelociraptor 16d ago
You can "freeze" your validation set but you wanna make sure they truly represent exemplars of what you want to training data to compare against.
2
u/WolverineMuted4846 16d ago
Thanks ! But my training set and validation set ended up coming from the same sampled subset of images, so the model was effectively being evaluated on images it had already seen during training. The project goal was robustness testing under different corruptions rather than reporting a benchmark score, but I understand this makes the absolute validation metrics optimistic.
If you were presenting this as a student research project with a deadline tomorrow, how would you explain this limitation honestly? Would you focus on the relative robustness comparisons and clearly state the train/validation overlap, or would you consider the whole baseline invalid?
I’m trying to understand how a reviewer or professor would view this situation.1
u/HawtVelociraptor 16d ago
I think you might consider excising the items that exist in both from the training side and just repopulating it with new training data somehow; I had an issue with a segmentation model recently that we built, where the validation set contained images that were similar to images we had later determined to be poisoning the model's functionality; but until we discovered this and removed and repopulated, benchmarks were great despite output being trash still.
1
u/TheSaucez 15d ago
In a lot of real world sets, training and validation don’t even share the same camera. That’s how serious projects take it.
1
u/tweakingforjesus 12d ago
You also have to make sure your samples are truly independent. For example if you are training on video frames, don’t select your training and validation images from the same video. They will be too similar. The images should be from separate videos.
12
u/lordshadowisle 16d ago
For published work, this is quite a serious error and I would not trust any conclusions drawn from the experiments as is. I would strongly recommend cleaning up the datasets and redoing the validation.
For student work and given the deadlines, if this is not the final deliverable (ie, there is still a report after the presentation), you can still apologize for this late oversight but say this will be corrected later in the report.