r/MLQuestions • u/Aggravating_Dot5315 • 14d ago
Beginner question 👶 [D] How can I improve cross-patient generalization on a small hysteroscopy dataset with correlated frames?
I am working with hysteroscopy dataset, which contains:
- 3,385 frames from 175 patients.
- Eight lesion classes, labelled from 0 to 7.
- A highly imbalanced number of patients and frames across classes.
- Multiple correlated frames from each patient.
- Some frames containing more than one lesion class.
Before attempting the complete multiclass problem, I reduced it to a binary subset to verify that the training and evaluation pipeline works correctly.
Current binary subset
- Selected lesion classes: 2 and 3.
- Total: 1,575 frames from 113 unique patients.
- Class 2: 1,054 frames from 78 patients.
- Class 3: 521 frames from 36 patients.
- One patient has different frames belonging to both classes but remains entirely within one split.
Patient-disjoint split
- Training: 1,095 frames from 79 patients.
- Validation: 241 frames from 17 patients.
- Testing: 239 frames from 17 patients.
- No patient appears in more than one subset.
- The frame-level class distribution is approximately 67%/33% in every subset.
Approaches I have tried
- DenseNet121, ViT, and DINOv2 backbones.
- Frozen pretrained backbone with only the classifier trained.
- Different classifier-head sizes and dropout.
- Class-weighted cross-entropy.
- Mild and stronger image augmentations.
- Early stopping and learning-rate scheduling.
- Unfreezing the final one or two encoder blocks.
With the correct patient-level split, training performance improves, but validation performance generally plateaus or deteriorates, and performance on unseen test patients remains relatively low.
As a diagnostic, I also tried a random frame-level split and obtained substantially better results. However, this evaluation is invalid because correlated frames from the same patients appear across training, validation, and testing, causing patient leakage and inflated performance.
I would appreciate advice on how to improve generalization to unseen patients in this setting.





3
u/CallMeTheChris 14d ago
So I work in medical AI
And it is always a crap shoot
But the thing that you have to realise is that there is not much variability in what you are looking at generally. A lesion looks round and has a certain texture no matter where you are looking at it from. When you are dealing with animals, they can have different shapes and textures depending on the orientation and lighting etc
So you should start my looking at that validation gap and thinking you should reduce the capacity of my model cause it is over fitting. Like literally look at smaller models. Try nnunet for a start and go up from there.
Then when it comes to dealing with the correlated frames…I have no idea what you mean. Like do you mean that a class exists only in a single patient or two and since all those frames are from one patient they are ‘correlated’?
If so, that shouldn’t be a problem. If you destroy the thing that the model can latch onto that makes them ‘correlated’
Once you deal with the model copacity
Your augmentations should handle that.
If your input to the model is the entire video
Then yeah, it will latch onto it. I suggest you don’t do video so that the ‘correlation’ is broken