Context
A free audio dataset of spoken digits. Think MNIST for audio. (3,000 recordings, 6 speakers )A simple audio/speech dataset consisting of recordings of spoken digits in wav files at 8kHz. The recordings are trimmed so that they have near minimal silence at the beginnings and ends.
FSDD is an open dataset, which means it will grow over time as data is contributed. In order to enable reproducibility and accurate citation the dataset is versioned using Zenodo DOI as well as git tags.
Current status6 speakers3,000 recordings (50 of each digit per speaker)English pronunciations
Created by:Zohar Jackson, César Souza, Jason Flaks, Yuxin Pan, Hereman Nicolas, & Adhish Thite.
Link:https://github.com/Jakobovski/free-spoken-digit-dataset
Content
What's inside is more than just rows and columns. Make it easy for others to get started by describing how you acquired the data and what time period it represents, too.
Acknowledgements
Zohar Jackson, César Souza, Jason Flaks, Yuxin Pan, Hereman Nicolas, & Adhish Thite. (2018, August 9). Jakobovski/free-spoken-digit-dataset: v1.0.8 (Version v1.0.8). Zenodo.http://doi.org/10.5281/zenodo.1342401
Inspiration
A free audio dataset of spoken digits. Think MNIST for audio.
