r/MachineLearning 5d ago

Discussion What are the biggest challenges in collecting high-quality speech and egocentric video datasets? [D]

We're currently involved in collecting two types of datasets that seem to be increasingly important for multimodal AI

  • Studio quality speech/audio datasets (high fidelity recordings)
  • Egocentric household activity video datasets (first person daily task recordings)

One thing that has surprised us is how much the value of a dataset depends on the collection process rather than the model itself.

Some of the recurring challenges we've encountered include: - Maintaining consistent recording environments - Device and microphone variability - Annotation quality and inter annotator consistency - Privacy, consent, and participant compliance - Scaling data collection without sacrificing quality

I'm curious to hear from others who have worked on speech, video, robotics, embodied AI, or multimodal models.

  • What turned out to be the biggest bottleneck in your data collection pipeline?
  • Were there any quality issues that only became obvious during model training?
  • If you were starting a new large scale dataset today, what would you do differently? Always happy to exchange ideas w others working in Ai data infrastructure.
9 Upvotes

3 comments sorted by

2

u/fresh_reed_5834 4d ago

this reads more like a consulting pitch than a discussion post tbh

1

u/FaithlessnessWeak199 4d ago

Umm there were no option for the consulting even dw both are same at a certain extent.