← Back to explorer

People's Speech

Type
dataset
Venue
MLCommons
Year
2026
Source
huggingface
Access
free
Language
en
Added
2026-07-17T19:59:20.671541+00:00
Verified
2026-07-17T19:59:20.671541+00:00

Summary

The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license.

Keywords

hf-dataset speech---asr crowdsourced machine-generated monolingual original parquet audio text datasets dask polars mlcroissant robust-speech-recognition noisy-speech-recognition speech-recognition has-paper

Topics

Speech / ASR

Research notes

  • downloads=41193; likes=275