•1 min read•from Machine Learning
What are the biggest challenges in collecting high-quality speech and egocentric video datasets? [D]
We're currently involved in collecting two types of datasets that seem to be increasingly important for multimodal AI
- Studio quality speech/audio datasets (high fidelity recordings)
- Egocentric household activity video datasets (first person daily task recordings)
One thing that has surprised us is how much the value of a dataset depends on the collection process rather than the model itself.
Some of the recurring challenges we've encountered include: - Maintaining consistent recording environments - Device and microphone variability - Annotation quality and inter annotator consistency - Privacy, consent, and participant compliance - Scaling data collection without sacrificing quality
I'm curious to hear from others who have worked on speech, video, robotics, embodied AI, or multimodal models.
- What turned out to be the biggest bottleneck in your data collection pipeline?
- Were there any quality issues that only became obvious during model training?
- If you were starting a new large scale dataset today, what would you do differently? Always happy to exchange ideas w others working in Ai data infrastructure.
[link] [comments]
Want to read more?
Check out the full article on the original site
Tagged with
#large dataset processing
#speech datasets
#video datasets
#egocentric video
#multimodal AI
#data collection
#annotation quality
#inter-annotator consistency
#privacy
#consent
#participant compliance
#recording environments
#device variability
#microphone variability
#high fidelity
#household activity
#data pipeline
#model training
#robotics
#embodied AI