datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ethiopian-languages-speech-dataset
Leyu Ethiopian Languages Speech Dataset
Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya.
Dataset Summary
Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti)
Total examples: 2750
License: CC-BY-4.0
Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/Benji-fish/ethiopian-languages-speech-dataset.fleurs-r-neucodec-all-languages
FLEURS-R NeuCodec All Languages
FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102
locales, plus a speaker label FLEURS itself does not ship.
Layout
data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the
viewer shows).
audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members
named audio/{locale}/{split}/{id}.wav (the path column).
neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.openslr-32-hq-SA-languages-Afrikaans
High quality TTS data for four South African languages - Afrikaans
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Afrikaans
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.benji-ethiopian-languages-speech-dataset
Leyu Ethiopian Languages Speech Dataset
Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya.
Dataset Summary
Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti)
Total examples: 2750
License: CC-BY-4.0
Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/LeyuCompetition/benji-ethiopian-languages-speech-dataset.openslr-32-hq-SA-languages
SLR32 – High Quality TTS Data for Four South African Languages
Identifier: SLR32License: CC BY-SA 4.0Source: https://www.openslr.org/32/
This dataset contains multi-speaker high quality transcribed audio data for four
languages of South Africa: Afrikaans (af_za), Sesotho (st_za),
Setswana (tn_za) and isiXhosa (xh_za).
The dataset consists of WAV files and a TSV file transcribing the audio.
In each folder the file line_index.tsv contains a FileID (which in turn
encodes the UserID)… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr-32-hq-SA-languages.speech-recognition-congolese-languages
Speech Recognition Datasets for Congolese Languages
Dataset Details
Dataset Description
This dataset contains two new benchmark corpora designed for low-resource languages spoken in the Democratic Republic of the Congo: The Lingala Read Speech Corpus LRSC, with 4.3 hours of labelled audio, and the Congolese Speech Radio Corpus CSRC, which offers 741 hours of unlabeled audio spanning four significant low-resource languages of the region (Lingala, Tshiluba… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/speech-recognition-congolese-languages.openslr-32-hq-SA-languages-Setswana
High quality TTS data for four South African languages - Setswana
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Setswana
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Setswana.openslr-32-hq-SA-languages-Sesotho
High quality TTS data for four South African languages - Sesotho
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Sesotho
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/openslr-32-hq-SA-languages-Sesotho.openslr-32-hq-SA-languages-Sesotho
High quality TTS data for four South African languages - Sesotho
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Sesotho
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Sesotho.openslr-32-hq-SA-languages-isiXhosa
High quality TTS data for four South African languages - isiXhosa
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - isiXhosa
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-isiXhosa.kasagadi
Kasagadi — Ghanaian Radio Broadcast Fact-Check Dataset
This is a multilingual dataset of transcribed, translated, and AI fact-checked segments from live radio broadcasts across Ghana. It is a Ghanaian initiative, covering Twi-language broadcasts from two Ghanaian FM stations and Hausa-language broadcasts from a third Ghanaian FM station serving Ghana's Zongo communities.
Dataset Summary
Station
Language
Broadcasts
Segments
Hours
Date Range
Angel FM
Twi… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/kasagadi.all-lab-speech
All Lab Speech
Cleaned African-language speech with embedded, playable audio (HF Audio). One config per language (<lang> = transcribed, <lang>_manifest = audio-only); the audio column sits right after audio_id and plays in the dataset viewer. Splits (train/validation/test) come from the source split labels.
from datasets import load_dataset
ds = load_dataset("African-Languages-Lab/all-lab-speech", "afrikaans")
Columns
audio_id, audio (playable), transcript… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/all-lab-speech.
