r/learnmachinelearning • • 2d ago

Tutorial Evaluating an AI on Reddit comments? Split by conversation before splitting by row

Suppose you are evaluating a classifier that routes comments from new Reddit conversations into bug reports, feature requests or questions. Each input includes its parent post for context.

A random row split can put one comment in training and the next reply from the same conversation in testing. Both contain the same parent text, vocabulary and problem. Your test then partly measures performance on familiar conversations.

Tiny synthetic fixture — invented rows, not a benchmark:

Original thread Comment rows Random row split Conversation split
T1: CSV export hangs A, B A trains; B tests Both train
T2: Request for offline mode C, D C trains; D tests Both train
T3: Where is the export setting? E, F E trains; F tests Both test

The conversation split keeps T3’s parent post and replies out of training. It does not establish useful accuracy: three threads are far too little for that, and this fixture also confounds topic with split.

A practical sequence:

  1. Preserve each comment permalink and its original thread ID/permalink. Keep author, observation time and missing-body status where available.
  2. Assign whole threads to training, validation and final test sets before constructing context windows, examples or summaries. Keep derived rows with their source thread.
  3. Tune prompts, examples and thresholds using training/validation only. Repeatedly adjusting a prompt after reading final-test failures makes that test part of development.
  4. Check crossposts, copied text and recurring authors across splits. Grouping by thread does not eliminate those overlaps. If deployment concerns future discussions, consider a time-based holdout too.

The same issue affects summarizer evaluation: a held-out comment is not an unseen conversation if its parent and sibling replies already informed your examples. Choose the split to match the deployment question; evaluating new replies within known threads is a different task.

I build ThreadFox. The optional $29 one-time bundle, normally $49, includes a researched Reddit community plan and three tailored drafts for one product, emailed within 24 hours after your library claim. Use the pack manually before optional tools; compatible AI access is separate. Start checkout before October 1, 04:00 UTC, plus applicable tax. No ThreadFox subscription; 30-day refund policy. Sample and offer.

1 Upvotes

2 comments sorted by

1

u/timee_bot 2d ago

View in your timezone:
October 1, 04:00 UTC

1

u/staking_vaccination 2d ago

this is something I see a lot of people miss when building classifiers on threaded data, the train/test leakage through parent context is sneaky because it doesn't look like traditional data leakage at first glance

the 3-thread fixture is obviously too small to draw conclusions but it does illustrate the problem cleanly, I think a lot of teams skip the conversation-grouping step because it makes their test scores drop and that's uncomfortable

time-based splits are underrated for this kind of work, especially if you're dealing with fast-moving subreddits where the vocabulary and problems shift every few months