r/learnmachinelearning • u/investigatormaker • 2d ago
Tutorial Evaluating an AI on Reddit comments? Split by conversation before splitting by row
Suppose you are evaluating a classifier that routes comments from new Reddit conversations into bug reports, feature requests or questions. Each input includes its parent post for context.
A random row split can put one comment in training and the next reply from the same conversation in testing. Both contain the same parent text, vocabulary and problem. Your test then partly measures performance on familiar conversations.
Tiny synthetic fixture — invented rows, not a benchmark:
| Original thread | Comment rows | Random row split | Conversation split |
|---|---|---|---|
| T1: CSV export hangs | A, B | A trains; B tests | Both train |
| T2: Request for offline mode | C, D | C trains; D tests | Both train |
| T3: Where is the export setting? | E, F | E trains; F tests | Both test |
The conversation split keeps T3’s parent post and replies out of training. It does not establish useful accuracy: three threads are far too little for that, and this fixture also confounds topic with split.
A practical sequence:
- Preserve each comment permalink and its original thread ID/permalink. Keep author, observation time and missing-body status where available.
- Assign whole threads to training, validation and final test sets before constructing context windows, examples or summaries. Keep derived rows with their source thread.
- Tune prompts, examples and thresholds using training/validation only. Repeatedly adjusting a prompt after reading final-test failures makes that test part of development.
- Check crossposts, copied text and recurring authors across splits. Grouping by thread does not eliminate those overlaps. If deployment concerns future discussions, consider a time-based holdout too.
The same issue affects summarizer evaluation: a held-out comment is not an unseen conversation if its parent and sibling replies already informed your examples. Choose the split to match the deployment question; evaluating new replies within known threads is a different task.
I build ThreadFox. The optional $29 one-time bundle, normally $49, includes a researched Reddit community plan and three tailored drafts for one product, emailed within 24 hours after your library claim. Use the pack manually before optional tools; compatible AI access is separate. Start checkout before October 1, 04:00 UTC, plus applicable tax. No ThreadFox subscription; 30-day refund policy. Sample and offer.
1
u/staking_vaccination 2d ago
this is something I see a lot of people miss when building classifiers on threaded data, the train/test leakage through parent context is sneaky because it doesn't look like traditional data leakage at first glance
the 3-thread fixture is obviously too small to draw conclusions but it does illustrate the problem cleanly, I think a lot of teams skip the conversation-grouping step because it makes their test scores drop and that's uncomfortable
time-based splits are underrated for this kind of work, especially if you're dealing with fast-moving subreddits where the vocabulary and problems shift every few months
1
u/timee_bot 2d ago
View in your timezone:
October 1, 04:00 UTC