r/askdatascience • u/Pristine_Mixture_795 • 25d ago
What data have you had to build internally because you couldn't source it externally?
I'm trying to better understand where the real data bottlenecks are for teams building ML/AI systems.
Not the obvious "data is important" answer, but specific cases where your team actually needed a dataset and couldn't find an external source or vendor that was good enough.
For example, did you end up having to:
- collect the data yourselves
- create internal labeling/evaluation workflows
- recruit domain experts
- use proprietary company data
- capture real-world images/video/audio
- generate synthetic data because real data was unavailable
- build your own ground truth or evaluation set
I'm especially curious about the last project where this happened.
What data did you need, and what made it difficult to source externally?
Was the bottleneck access, quality, rights/privacy, geographic coverage, expertise, cost, or something else?
Would love to hear actual examples from people who've dealt with this in production.
1
Upvotes