voicedata/pidginData
Naija-ASR-Corpus v2.0 (NAC-v2.0) A Foundational Automatic Speech Recognition Corpus for Nigerian Pidgin (Naija, PCM) 📌 Dataset Summary Naija-ASR-Corpus (NAC-v2.0) is a speech dataset derived from the Universal Dependencies Naija Spoken Corpus (UD_Naija-NSC). The NAC Team processed the original long-form recordings by: Segmenting the audio into sentence-level clips. Aligning each clip to its transcript (text_ortho) from the CoNLL-U source. Tagging each sample… See the full description on the dataset page: https://huggingface.co/datasets/voicedata/pidginData.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face