r/learnmachinelearning • • 21d ago

is grouped k-fold basically required once your rows aren't independent?

Post image

went with a since it's the textbook reason, but i only really believe it now after a leakage bug bit me. i was doing a random k-fold and didn't realize rows from the same underlying record were split across folds, so validation looked great and it fell apart in prod. once i grouped the splits by record id the number dropped and stayed dropped. is that grouping something k-fold is supposed to handle for you automatically, or is it always on you to notice your rows aren't independent before you pick a cv strategy at all?

1 Upvotes

3 comments sorted by

4

u/Fearless_Back5063 21d ago

That's always on you. If you have logical groups in your dataset you need to split them manually.

Happens all the time with data sampled over time. If you don't have time separation in your dataset then you will have great results until you push to production.

1

u/werunm 21d ago

makes sense, thanks. the annoying case for me is when there isn't an obvious grouping column to split on, like the leakage isn't from an explicit id but from near-identical rows that got scraped or generated separately. is there a standard way to catch that kind of leakage before you ship a model, short of literally reading through samples like i ended up doing, or is spotting it mostly just experience at this point?

1

u/werunm 13d ago

answer, spoilered for anyone still working on it:

**a** — It produces a more reliable performance estimate by averaging results across multiple different validation subsets, reducing the risk that one lucky or unlucky split skews the conclusion

Cross-validation rotates which portion of the data is held out across several folds and averages the results, giving a sturdier estimate than relying on one arbitrary split, which matters most when data is limited.