Omnilingual ASR Corpus
- Type
- dataset
- Venue
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- aae, aal, aao, abn, abr, abs, abv, acm, acw, acx, adf, aeb, aec, afb, afo, ahl, ahs, aju, ala, aln, alo, amu, anc, ank, anp, anw, aom, apc, apd, arq, ars, ary, arz, avl, awo, ayl, ayp, bbu, bcs, bcy, bda, bde, bdm, bew, bhb, bhh, bho, bhp, bhr, bjj, bjk, bjn, bjt, bky, bmm, bmq, bns, bou, bqg, bra, brh, brx, bsj, btm, bug, buo, bux, bwr, bxf, byc, bys, byx, bzc, bzw, ccg, cen, cfa, cgg, chq, ckl, ckr, cky, cte, ctl, dbd, dcc, deg, dgh, dje, dty, dzg, ebu, ego, eiv, ekr, elm, ets, etu, ext, eyo, fat, ffm, fia, fip, fkk, fuc, fue, fuf, fuh, fui, fuq, fuv, gbm, gbr, gby, gcc, gdf, ges, gjk, glw, gol, gom, gsl, gui, gur, guz, gwe, gyz, hah, hao, haw, hbb, hz, hia, hkk, hla, hno, hoj, hue, hul, hwo, ida, idu, ijc, ijn, ikw, ish, iso, its, itw, itz, jal, jax, jmx, jns, juk, juo, kai, kaj, ks, kbl, kbt, kcq, keu, kfe, kfk, kfp, kjc, kjk, kmy, kna, knn, kol, koo, kpo, kqo, ksd, kto, kj, kuh, kwm, kxp, kyx, lag, lcm, ldb, lij, lir, lkb, lla, lnu, loa, lto, lus, lwg, mab, maf, max, mde, mek, mer, meu, mfm, mfn, mfo, mfv, mgi, mig, miu, mkf, mlq, mne, mqy, mrr, mrt, msh, msw, mtr, mtu, mtx, mui, mxs, mxy, mzl, nal, nap, nbh, ncf, nco, ndi, ng, ngi, nhg, nhn, nhq, nja, noe, odk, odu, ogo, orc, pbs, pbt, pbu, pex, phr, pip, piy, pko, plt, pmq, pms, pmy, pnb, poc, poe, pow, pst, qug, qum, quv, rag, rob, rof, roo, rth, sau, say, scn, shu, si, sip, siw, sjr, skg, snc, snk, sol, sps, src, sro, ste, sua, tan, tbf, tcf, tcy, tdn, tdx, tgc, the, thq, thr, thv, tio, tkg, tkt, tlp, tpl, tpz, tqp, trp, trq, ttj, ttr, ttu, tul, tuq, tuv, tuy, tvo, twu, txs, txy, uki, uzn, vai, ver, vmc, vmj, vmm, vmp, vmz, vro, wci, weo, wja, wji, wof, xmv, xmw, xpe, xti, xtu, yay, ydd, yer, yes, zga, zoh, zor, zpv, zpy, ztg, ztn, ztp, zts, ztu
- Added
- 2026-07-17T19:59:08.976683+00:00
- Verified
- 2026-07-17T19:59:08.976683+00:00
Summary
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project ([blog](https://ai.meta.com/blog/omnilingual-asr-advancing-automatic-speech-recognition/), [model](https://github.com/facebookresearch/omnilingual-asr), [paper](https://ai.meta.com/research/publications/omnilingual-asr-open-source-multilingual-speech-recognition-for-1600-languages/)) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Keywords
hf-dataset speech---asr parquet optimized-parquet audio text datasets dask polars mlcroissant has-paper
Topics
Speech / ASR
Research notes
- downloads=14033; likes=207