CoolFace
Datasetpublic

Lindo20/lwazi-asr-corpus-compressed

Lwazi ASR Corpus Collection This repository contains a curated collection of the Lwazi Automatic Speech Recognition (ASR) Corpus for several low-resourced South African languages. These datasets are designed for use in speech recognition research and development, particularly for underrepresented languages. Corpus Overview Each corpus consists of scripted telephonic speech recordings collected from native speakers, along with corresponding transcriptions. The… See the full description on the dataset page: https://huggingface.co/datasets/Lindo20/lwazi-asr-corpus-compressed.

sourceHugging Facecc-by-3.0updated 7mo agoView on Hugging Face
0likes10downloads
Dataset Card

Lwazi ASR Corpus Collection

This repository contains a curated collection of the Lwazi Automatic Speech Recognition (ASR) Corpus for several low-resourced South African languages. These datasets are designed for use in speech recognition research and development, particularly for underrepresented languages.

Corpus Overview

Each corpus consists of scripted telephonic speech recordings collected from native speakers, along with corresponding transcriptions. The audio recordings are primarily from general domain telephone conversations. All data is licensed under Creative Commons BY 2.5, allowing for open research and non-commercial use with attribution.

LanguageScriptedLicenseSpeakersHoursUtterancesDomain
isiNdebele✅CC BY 3.0102006,013Telephonic/General
isiXhosa✅CC BY 3.092106,242Telephonic/General
isiZulu✅CC BY 3.081995,785Telephonic/General
Sepedi✅CC BY 3.091905,640Telephonic/General
Sesotho✅CC BY 3.072026,027Telephonic/General
Setswana✅CC BY 3.082035,970Telephonic/General
Siswati✅CC BY 3.0101965,838Telephonic/General
Tshivenda✅CC BY 3.071985,939Telephonic/General
Xitsonga✅CC BY 3.082146,426Telephonic/General

Use Cases

This corpus is particularly useful for:

  • —Training and evaluating ASR models for low-resourced South African languages.
  • —Linguistic analysis and phonetic research.
  • —Data augmentation and transfer learning tasks in multilingual NLP/ASR.

Data Structure

Each language corpus includes:

  • —audio/: Subfolders per speaker that contain WAV audio files sampled from telephonic speech.
  • —transcriptions/: Text transcripts aligned with audio files.
  • —metadata.csv: Speaker information, durations, and utterance IDs.

Licensing

This dataset is licensed under Creative Commons Attribution 3.0 (CC BY 3.0). You are free to share and adapt the data with appropriate attribution.The dictionaries made availabe on this site are derived works of the "NCHLT-inlang Pronunciation Dictionaries" by the Meraka Institute, CSIR and the North-West University, available from the RMA and released under a Creative Commons Attribution 3.0 Unported License (CC BY 3.0). When using these dictionaries, please cite the following papers:

Citation

E. Barnard, M. Davel and C. van Heerden, "ASR Corpus Design for Resource-Scarce Languages," in Proceedings of the 10th Annual Conference of the International Speech Communication Association (Interspeech), Brighton, United Kingdom, September 2009, pp. 2847-2850.