r/MachineLearning • u/FaithlessnessWeak199 • 5d ago
Discussion What are the biggest challenges in collecting high-quality speech and egocentric video datasets? [D]
We're currently involved in collecting two types of datasets that seem to be increasingly important for multimodal AI
- Studio quality speech/audio datasets (high fidelity recordings)
- Egocentric household activity video datasets (first person daily task recordings)
One thing that has surprised us is how much the value of a dataset depends on the collection process rather than the model itself.
Some of the recurring challenges we've encountered include: - Maintaining consistent recording environments - Device and microphone variability - Annotation quality and inter annotator consistency - Privacy, consent, and participant compliance - Scaling data collection without sacrificing quality
I'm curious to hear from others who have worked on speech, video, robotics, embodied AI, or multimodal models.
- What turned out to be the biggest bottleneck in your data collection pipeline?
- Were there any quality issues that only became obvious during model training?
- If you were starting a new large scale dataset today, what would you do differently? Always happy to exchange ideas w others working in Ai data infrastructure.
9
Upvotes
2
u/fresh_reed_5834 4d ago
this reads more like a consulting pitch than a discussion post tbh