CoolFace
Datasetpublic

Harcuracy/google_waxal_asr_challenge

WaxalNLP ASR — Cleaned Subset (Lingala, Shona, Luganda) This dataset is a cleaned, corrected subset of google/WaxalNLP, covering the train and validation splits for three languages: lin_asr — Lingala sna_asr — Shona lug_asr — Luganda The test split from the original dataset is intentionally excluded. What was changed The original transcriptions for these three languages contained a number of errors. A corrected transcription file was applied on top of the… See the full description on the dataset page: https://huggingface.co/datasets/Harcuracy/google_waxal_asr_challenge.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes195downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Harcuracy/google_waxal_asr_challenge · CoolFace