← Back to explorer

Omnilingual ASR Corpus

Type
dataset
Venue
facebook
Year
2026
Source
huggingface
Access
free
Language
aae, aal, aao, abn, abr, abs, abv, acm, acw, acx, adf, aeb, aec, afb, afo, ahl, ahs, aju, ala, aln, alo, amu, anc, ank, anp, anw, aom, apc, apd, arq, ars, ary, arz, avl, awo, ayl, ayp, bbu, bcs, bcy, bda, bde, bdm, bew, bhb, bhh, bho, bhp, bhr, bjj, bjk, bjn, bjt, bky, bmm, bmq, bns, bou, bqg, bra, brh, brx, bsj, btm, bug, buo, bux, bwr, bxf, byc, bys, byx, bzc, bzw, ccg, cen, cfa, cgg, chq, ckl, ckr, cky, cte, ctl, dbd, dcc, deg, dgh, dje, dty, dzg, ebu, ego, eiv, ekr, elm, ets, etu, ext, eyo, fat, ffm, fia, fip, fkk, fuc, fue, fuf, fuh, fui, fuq, fuv, gbm, gbr, gby, gcc, gdf, ges, gjk, glw, gol, gom, gsl, gui, gur, guz, gwe, gyz, hah, hao, haw, hbb, hz, hia, hkk, hla, hno, hoj, hue, hul, hwo, ida, idu, ijc, ijn, ikw, ish, iso, its, itw, itz, jal, jax, jmx, jns, juk, juo, kai, kaj, ks, kbl, kbt, kcq, keu, kfe, kfk, kfp, kjc, kjk, kmy, kna, knn, kol, koo, kpo, kqo, ksd, kto, kj, kuh, kwm, kxp, kyx, lag, lcm, ldb, lij, lir, lkb, lla, lnu, loa, lto, lus, lwg, mab, maf, max, mde, mek, mer, meu, mfm, mfn, mfo, mfv, mgi, mig, miu, mkf, mlq, mne, mqy, mrr, mrt, msh, msw, mtr, mtu, mtx, mui, mxs, mxy, mzl, nal, nap, nbh, ncf, nco, ndi, ng, ngi, nhg, nhn, nhq, nja, noe, odk, odu, ogo, orc, pbs, pbt, pbu, pex, phr, pip, piy, pko, plt, pmq, pms, pmy, pnb, poc, poe, pow, pst, qug, qum, quv, rag, rob, rof, roo, rth, sau, say, scn, shu, si, sip, siw, sjr, skg, snc, snk, sol, sps, src, sro, ste, sua, tan, tbf, tcf, tcy, tdn, tdx, tgc, the, thq, thr, thv, tio, tkg, tkt, tlp, tpl, tpz, tqp, trp, trq, ttj, ttr, ttu, tul, tuq, tuv, tuy, tvo, twu, txs, txy, uki, uzn, vai, ver, vmc, vmj, vmm, vmp, vmz, vro, wci, weo, wja, wji, wof, xmv, xmw, xpe, xti, xtu, yay, ydd, yer, yes, zga, zoh, zor, zpv, zpy, ztg, ztn, ztp, zts, ztu
Added
2026-07-17T19:59:08.976683+00:00
Verified
2026-07-17T19:59:08.976683+00:00

Summary

The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project ([blog](https://ai.meta.com/blog/omnilingual-asr-advancing-automatic-speech-recognition/), [model](https://github.com/facebookresearch/omnilingual-asr), [paper](https://ai.meta.com/research/publications/omnilingual-asr-open-source-multilingual-speech-recognition-for-1600-languages/)) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.

Keywords

hf-dataset speech---asr parquet optimized-parquet audio text datasets dask polars mlcroissant has-paper

Topics

Speech / ASR

Research notes

  • downloads=14033; likes=207