KasuleTrevor/Lingala_100hrs
license: cc-by-4.0 language: - ln task_categories: - automatic-speech-recognition pretty_name: Lingala 100hrs ASR Lingala 100hrs 110.7 hours (23,539 rows) of Lingala speech with transcriptions, aggregated from three publicly available CC-BY-4.0 corpora for ASR research. Composition Counts from a full-pass audit on 2026-07-09: Source Upstream location Rows Splits AfriVoice (Lingala)… See the full description on the dataset page: https://huggingface.co/datasets/KasuleTrevor/Lingala_100hrs.
license: cc-by-4.0 language:
- ln task_categories:
- automatic-speech-recognition pretty_name: Lingala 100hrs ASR ---
Lingala 100hrs
110.7 hours (23,539 rows) of Lingala speech with transcriptions, aggregated from three publicly available CC-BY-4.0 corpora for ASR research.
Composition
Counts from a full-pass audit on 2026-07-09:
Splits
Licensing
The three documented sources are all CC-BY-4.0, so this aggregation is distributed under CC-BY-4.0 with the attributions below.
Dataset structure
audio: audio (mono, 16 kHz) + sampling ratetext: Lingala transcriptsource: upstream corpus for the row (Afrivoice,LRSC,fleurs)
Citations
LRSC:
Kimanuka, U., wa Maina, C., & Büyük, O. (2023). Speech Recognition Datasets for Low-resource Congolese Languages. AfricaNLP workshop at ICLR 2023. Data: Mendeley Data, V1, doi:10.17632/28x8tc9n9k.1
FLEURS:
Conneau, A., et al. (2022). FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech. IEEE SLT 2022.
AfriVoice:
Digital Umuganda. AfriVoice: multilingual African speech dataset (Shona, Lingala, Fulani, Malagasy, Wolof, Somali). https://huggingface.co/datasets/DigitalUmuganda/AfriVoice
