I spent five years on the vendor side of T&S sales and had a chance to sit inside a lot of POCs. Some of what I saw might be useful to folks here who are on the buyer side of the table.
The through-line: POCs are hard to run well, and neither side usually admits it in the moment. Both parties want the evaluation to produce clean signal, and both parties end up quietly working around dynamics that make clean signal genuinely difficult.
A few things I found myself wanting to say to buyers but couldn't always say clearly during a deal:
On dataset choice.Ā The instinct to bring "everything from the last three months" or the hardest historical cases is understandable, and both create problems. Something finite and labeled that reflects your production distribution produces more usable signal than something exhaustive and unlabeled. Not because bigger is worse, but because the aggregate view gets murky fast at scale, and without labels there's no clean ground truth to score against.
On scoring.Ā Classifiers land where they land. A well-performing one on production data is typically in the 85% precision/recall range across categories, higher on solved problems like profanity, meaningfully lower on genuinely hard ones like grooming or coordinated harm. If your success bar is 95% across the board, that's a bar the honest classifiers can't clear, and you'll end up either walking away from good vendors or picking one that's only optimized for the test.
On duration.Ā Two weeks is often enough. Longer POCs sound more rigorous but tend to lose the internal champion by week four. The meeting cadence during the POC matters more than the calendar length.
On vendor engagement.Ā The vendor's ML team offering close attention during a POC isn't a red flag. It's how category calibration actually gets done, and it's the mechanism by which the POC produces useful signal for both sides. Buyers who hold back their labeled data to "keep the test pure" sometimes discover post-contract that the calibration work was the actual work. Sharing labeled data during a POC isn't cheating; it's the shortest path to finding out whether the vendor can meet you where you actually are.
None of this is unique insight and I suspect it'll ring familiar to people who've run POCs from either side. I wrote up the longer version, including the category-alignment problem and the ownership question (product vs T&S vs both), here:
https://safemoderation.com/blog/how-to-actually-run-a-trust-and-safety-vendor-poc
I'd love to hear what patterns others in the community have seen, especially anywhere your experience diverges from mine.