CoolFace
Datasetpublic

KasuleTrevor/Lingala_100hrs

license: cc-by-4.0 language: - ln task_categories: - automatic-speech-recognition pretty_name: Lingala 100hrs ASR Lingala 100hrs 110.7 hours (23,539 rows) of Lingala speech with transcriptions, aggregated from three publicly available CC-BY-4.0 corpora for ASR research. Composition Counts from a full-pass audit on 2026-07-09: Source Upstream location Rows Splits AfriVoice (Lingala)… See the full description on the dataset page: https://huggingface.co/datasets/KasuleTrevor/Lingala_100hrs.

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes267downloads
Dataset Card

license: cc-by-4.0 language:

  • ln task_categories:
  • automatic-speech-recognition pretty_name: Lingala 100hrs ASR ---

Lingala 100hrs

110.7 hours (23,539 rows) of Lingala speech with transcriptions, aggregated from three publicly available CC-BY-4.0 corpora for ASR research.

Composition

Counts from a full-pass audit on 2026-07-09:

SourceUpstream locationRowsSplits
AfriVoice (Lingala)https://huggingface.co/datasets/DigitalUmuganda/AfriVoice17,544train (16,144), validation (915), test (485)
LRSC (Lingala Read Speech Corpus)https://data.mendeley.com/datasets/28x8tc9n9k/12,937train (2,489), validation (65), test (383)
FLEURS (google/fleurs, config ln_cd)https://huggingface.co/datasets/google/fleurs2,995train (2,526), validation (65), test (404)

Splits

SplitRowsHoursMean durationMean words
train21,159100.017.0 s24.1
validation1,0455.117.6 s24.9
test1,3355.615.1 s19.9

Licensing

The three documented sources are all CC-BY-4.0, so this aggregation is distributed under CC-BY-4.0 with the attributions below.

SourceLicenseVerified against upstream on
AfriVoice (Lingala)CC-BY-4.0 (upstream repo is gated: users must accept its access terms)2026-07-09
LRSCCC BY 4.0 (Mendeley Data, DOI 10.17632/28x8tc9n9k.1)2026-07-09
FLEURSCC-BY-4.02026-07-09

Dataset structure

  • audio: audio (mono, 16 kHz) + sampling rate
  • text: Lingala transcript
  • source: upstream corpus for the row (Afrivoice, LRSC, fleurs)

Citations

LRSC:

Kimanuka, U., wa Maina, C., & Büyük, O. (2023). Speech Recognition Datasets for Low-resource Congolese Languages. AfricaNLP workshop at ICLR 2023. Data: Mendeley Data, V1, doi:10.17632/28x8tc9n9k.1

FLEURS:

Conneau, A., et al. (2022). FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech. IEEE SLT 2022.

AfriVoice:

Digital Umuganda. AfriVoice: multilingual African speech dataset (Shona, Lingala, Fulani, Malagasy, Wolof, Somali). https://huggingface.co/datasets/DigitalUmuganda/AfriVoice