1 min readfrom Machine Learning

What are the biggest challenges in collecting high-quality speech and egocentric video datasets? [D]

We're currently involved in collecting two types of datasets that seem to be increasingly important for multimodal AI

  • Studio quality speech/audio datasets (high fidelity recordings)
  • Egocentric household activity video datasets (first person daily task recordings)

One thing that has surprised us is how much the value of a dataset depends on the collection process rather than the model itself.

Some of the recurring challenges we've encountered include: - Maintaining consistent recording environments - Device and microphone variability - Annotation quality and inter annotator consistency - Privacy, consent, and participant compliance - Scaling data collection without sacrificing quality

I'm curious to hear from others who have worked on speech, video, robotics, embodied AI, or multimodal models.

  • What turned out to be the biggest bottleneck in your data collection pipeline?
  • Were there any quality issues that only became obvious during model training?
  • If you were starting a new large scale dataset today, what would you do differently? Always happy to exchange ideas w others working in Ai data infrastructure.
submitted by /u/FaithlessnessWeak199
[link] [comments]

Want to read more?

Check out the full article on the original site

View original article

Tagged with

#large dataset processing
#speech datasets
#video datasets
#egocentric video
#multimodal AI
#data collection
#annotation quality
#inter-annotator consistency
#privacy
#consent
#participant compliance
#recording environments
#device variability
#microphone variability
#high fidelity
#household activity
#data pipeline
#model training
#robotics
#embodied AI