People's Speech
- Type
- dataset
- Venue
- MLCommons
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en
- Added
- 2026-07-17T19:59:20.671541+00:00
- Verified
- 2026-07-17T19:59:20.671541+00:00
Summary
The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license.
Keywords
hf-dataset speech---asr crowdsourced machine-generated monolingual original parquet audio text datasets dask polars mlcroissant robust-speech-recognition noisy-speech-recognition speech-recognition has-paper
Topics
Speech / ASR
Research notes
- downloads=41193; likes=275