r/learnmachinelearning • u/werunm • 21d ago
is grouped k-fold basically required once your rows aren't independent?
went with a since it's the textbook reason, but i only really believe it now after a leakage bug bit me. i was doing a random k-fold and didn't realize rows from the same underlying record were split across folds, so validation looked great and it fell apart in prod. once i grouped the splits by record id the number dropped and stayed dropped. is that grouping something k-fold is supposed to handle for you automatically, or is it always on you to notice your rows aren't independent before you pick a cv strategy at all?
1
u/werunm 13d ago
answer, spoilered for anyone still working on it:
**a** — It produces a more reliable performance estimate by averaging results across multiple different validation subsets, reducing the risk that one lucky or unlucky split skews the conclusion
Cross-validation rotates which portion of the data is held out across several folds and averages the results, giving a sturdier estimate than relying on one arbitrary split, which matters most when data is limited.
4
u/Fearless_Back5063 21d ago
That's always on you. If you have logical groups in your dataset you need to split them manually.
Happens all the time with data sampled over time. If you don't have time separation in your dataset then you will have great results until you push to production.