datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ethiopian-languages-speech-dataset
Leyu Ethiopian Languages Speech Dataset
Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya.
Dataset Summary
Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti)
Total examples: 2750
License: CC-BY-4.0
Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/Benji-fish/ethiopian-languages-speech-dataset.fleurs-r-neucodec-all-languages
FLEURS-R NeuCodec All Languages
FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102
locales, plus a speaker label FLEURS itself does not ship.
Layout
data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the
viewer shows).
audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members
named audio/{locale}/{split}/{id}.wav (the path column).
neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.openslr-32-hq-SA-languages-Afrikaans
High quality TTS data for four South African languages - Afrikaans
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Afrikaans
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.benji-ethiopian-languages-speech-dataset
Leyu Ethiopian Languages Speech Dataset
Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya.
Dataset Summary
Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti)
Total examples: 2750
License: CC-BY-4.0
Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/LeyuCompetition/benji-ethiopian-languages-speech-dataset.openslr-32-hq-SA-languages
SLR32 – High Quality TTS Data for Four South African Languages
Identifier: SLR32License: CC BY-SA 4.0Source: https://www.openslr.org/32/
This dataset contains multi-speaker high quality transcribed audio data for four
languages of South Africa: Afrikaans (af_za), Sesotho (st_za),
Setswana (tn_za) and isiXhosa (xh_za).
The dataset consists of WAV files and a TSV file transcribing the audio.
In each folder the file line_index.tsv contains a FileID (which in turn
encodes the UserID)… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr-32-hq-SA-languages.speech-recognition-congolese-languages
Speech Recognition Datasets for Congolese Languages
Dataset Details
Dataset Description
This dataset contains two new benchmark corpora designed for low-resource languages spoken in the Democratic Republic of the Congo: The Lingala Read Speech Corpus LRSC, with 4.3 hours of labelled audio, and the Congolese Speech Radio Corpus CSRC, which offers 741 hours of unlabeled audio spanning four significant low-resource languages of the region (Lingala, Tshiluba… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/speech-recognition-congolese-languages.openslr-32-hq-SA-languages-Setswana
High quality TTS data for four South African languages - Setswana
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Setswana
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Setswana.openslr-32-hq-SA-languages-Sesotho
High quality TTS data for four South African languages - Sesotho
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Sesotho
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/openslr-32-hq-SA-languages-Sesotho.whisper-eval-rare-languages-csv
Whisper 3 Large Evaluation on Mozilla Common Voice 17 Rare Languages (Enhanced Metrics)
Dataset Description
This enhanced dataset contains comprehensive evaluation results of OpenAI's Whisper 3 Large model on rare languages from Mozilla Common Voice 17, with extensive additional metrics for thorough ASR evaluation.
Key Features
Enhanced Error Metrics:
WER (Word Error Rate): Standard word-level error measurement
CER (Character Error Rate): Character-level error… See the full description on the dataset page: https://huggingface.co/datasets/norbertm/whisper-eval-rare-languages-csv.openslr-32-hq-SA-languages-Sesotho
High quality TTS data for four South African languages - Sesotho
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Sesotho
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Sesotho.openslr-32-hq-SA-languages-isiXhosa
High quality TTS data for four South African languages - isiXhosa
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - isiXhosa
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-isiXhosa.17-minute-world-languages_allemande
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/allemande/
Site à scrapper
17-minute-world-languages_georgien
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/géorgienne/
Site à scrapper
17-minute-world-languages_malais
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/malaise/
Site à scrapper
17-minute-world-languages_jordanien
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/jordanienne/
Site à scrapper
17-minute-world-languages_malgache
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/malgache/
Site à scrapper
17-minute-world-languages_persan
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/persane/
Site à scrapper
kasagadi
Kasagadi — Ghanaian Radio Broadcast Fact-Check Dataset
This is a multilingual dataset of transcribed, translated, and AI fact-checked segments from live radio broadcasts across Ghana. It is a Ghanaian initiative, covering Twi-language broadcasts from two Ghanaian FM stations and Hausa-language broadcasts from a third Ghanaian FM station serving Ghana's Zongo communities.
Dataset Summary
Station
Language
Broadcasts
Segments
Hours
Date Range
Angel FM
Twi… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/kasagadi.17-minute-world-languages_bengali
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/bengali/
Site à scrapper
17-minute-world-languages_mongol
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/mongole/
Site à scrapper
17-minute-world-languages_amharique
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/amharique/
Site à scrapper
17-minute-world-languages_afrikaans
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/afrikaans/
Site à scrapper
17-minute-world-languages_albanaise
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/albanaise/
Site à scrapper
17-minute-world-languages_azeri
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/azéri/
Site à scrapper
17-minute-world-languages_bosniaque
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/bosniaque/
Site à scrapper
17-minute-world-languages_dari
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/dari/
Site à scrapper
17-minute-world-languages_egyptien
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/égyptienne/
Site à scrapper
17-minute-world-languages_lingala
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/lingala/
Site à scrapper
17-minute-world-languages_slovene
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/slovène/
Site à scrapper
17-minute-world-languages_swahili
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/swahilie/
Site à scrapper
