CoolFace
Datasetpublic

lindonghello/omnilingual-asr-corpus

Meta Omnilingual ASR Corpus The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models. Data schema { `language`: "lij_Latn", `iso_639_3`: "lij", `iso_15924`: "Latn"… See the full description on the dataset page: https://huggingface.co/datasets/lindonghello/omnilingual-asr-corpus.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes113downloads
settings

This repository belongs to lindonghello on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameomnilingual-asr-corpus
visibilitypublic
licencecc-by-4.0
gatedno
ownerlindonghello
Account settings
lindonghello/omnilingual-asr-corpus · CoolFace