datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multilingual-TTS-language
Multilingual-TTS-language
malaysia-ai/Multilingual-TTS
with two extra columns:
column
description
audio_filename, text, speaker
unchanged from malaysia-ai/Multilingual-TTS
language
language detected from the text (transcript) column of every row
post-normalized
text after rule-based punctuation / capitalization normalization
All original columns and the file/folder layout are preserved: 1552 subsets /
1561 parquet files, 121,842,371 rows (August 2026). Every… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS-language.ethiopian-languages-speech-dataset
Leyu Ethiopian Languages Speech Dataset
Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya.
Dataset Summary
Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti)
Total examples: 2750
License: CC-BY-4.0
Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/Benji-fish/ethiopian-languages-speech-dataset.samromur_childrenThe Samrómur Children corpus contains more than 137000 validated speech-recordings uttered by Icelandic children.ePark_zu_yu_duan_wen_indigenous_language_essays
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_zu_yu_duan_wen_indigenous_language_essays
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_zu_yu_duan_wen_indigenous_language_essays.Multimodal_Greek_Sign_Language_and_Lip_Reading
Multimodal_Greek_Sign_Language_and_Lip_Reading (v1)
The Multimodal_Greek_Sign_Language_and_Lip_Reading (v1) is a comprehensive dataset designed for research and development in multimodal machine learning, speech recognition, vision recognition, sign language recognition, sign language translation and accessibility technologies.
Description
The dkourem/Multimodal_Greek_Sign_Language_and_Lip_Reading-v1 is a comprehensive dataset designed for research and development… See the full description on the dataset page: https://huggingface.co/datasets/dkourem/Multimodal_Greek_Sign_Language_and_Lip_Reading.fleurs-r-neucodec-all-languages
FLEURS-R NeuCodec All Languages
FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102
locales, plus a speaker label FLEURS itself does not ship.
Layout
data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the
viewer shows).
audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members
named audio/{locale}/{split}/{id}.wav (the path column).
neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.samromur_milljonSamrómur Milljón consists of approximately 1 million of speech recordings (967 hours) collected through the platform samromur.is; the transcripts accompanying these recordings were automatically verified using various ASR systems such as: Wav2Vec, Whisper and NeMo.raddromur_asr
Dataset Card for raddromur_asr
Dataset Summary
The "Raddrómur Icelandic Speech 22.09" ("Raddrómur Corpus" for short) is an Icelandic corpus created by the Language and Voice Laboratory (LVL) at Reykjavík University (RU) in 2022. It is made out of radio podcasts mostly taken from RÚV (ruv.is).
Example Usage
The Raddrómur Corpus counts with the train split only. To load the training split pass its name as a config name:
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/language-and-voice-lab/raddromur_asr.samromur_asrSamrómur Icelandic Speech 1.0.ePark_qing_jing_zu_yu_contextual_indigenous_language
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_qing_jing_zu_yu_contextual_indigenous_language
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_qing_jing_zu_yu_contextual_indigenous_language.althingi_asrAlthingi Parliamentary Speech consists of approximately 542 hours of recorded speech from Althingi, the Icelandic Parliament. Speeches date from 2005-2016.shan_language_asr_voices
⭐ A Voice for the Shan People: The SHAN Herald Agency Audio Archive
This is an extensive 306-hour audio dataset of the Shan (Tai-Yai) language, meticulously curated from the public broadcasts of the Shan Herald Agency for News (SHAN). For over two decades, SHAN has been a vital, independent voice for the people of Shan State, Myanmar, chronicling their stories of culture, politics, and the enduring struggle for federal democracy.
This archive stands as one of the largest publicly… See the full description on the dataset page: https://huggingface.co/datasets/freococo/shan_language_asr_voices.openslr-32-hq-SA-languages-Afrikaans
High quality TTS data for four South African languages - Afrikaans
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Afrikaans
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.malromur_asrThe Málrómur corpus is an open source corpus of Icelandic voice samples.benji-ethiopian-languages-speech-dataset
Leyu Ethiopian Languages Speech Dataset
Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya.
Dataset Summary
Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti)
Total examples: 2750
License: CC-BY-4.0
Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/LeyuCompetition/benji-ethiopian-languages-speech-dataset.Afar-language-text-to-speech-TTS
Usage
This dataset is designed to support the development of Text-to-Speech (TTS) systems for the Afar language. It can be integrated into web applications, mobile apps, desktop software, or other platforms that require natural-sounding Afar voice synthesis or accurate spoken language recognition.
For applications involving virtual avatars or voice personas, the following culturally appropriate voice names are recommended:
Female Voices: Emeli, Hanaawi, Kareera, Laysani, Kulsuma… See the full description on the dataset page: https://huggingface.co/datasets/Charif-Ayfarah/Afar-language-text-to-speech-TTS.shan_language_asr_voices
⭐ A Voice for the Shan People: The SHAN Herald Agency Audio Archive
This is an extensive 306-hour audio dataset of the Shan (Tai-Yai) language, meticulously curated from the public broadcasts of the Shan Herald Agency for News (SHAN). For over two decades, SHAN has been a vital, independent voice for the people of Shan State, Myanmar, chronicling their stories of culture, politics, and the enduring struggle for federal democracy.
This archive stands as one of the largest… See the full description on the dataset page: https://huggingface.co/datasets/myandev/shan_language_asr_voices.openslr-32-hq-SA-languages
SLR32 – High Quality TTS Data for Four South African Languages
Identifier: SLR32License: CC BY-SA 4.0Source: https://www.openslr.org/32/
This dataset contains multi-speaker high quality transcribed audio data for four
languages of South Africa: Afrikaans (af_za), Sesotho (st_za),
Setswana (tn_za) and isiXhosa (xh_za).
The dataset consists of WAV files and a TSV file transcribing the audio.
In each folder the file line_index.tsv contains a FileID (which in turn
encodes the UserID)… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr-32-hq-SA-languages.speech-recognition-congolese-languages
Speech Recognition Datasets for Congolese Languages
Dataset Details
Dataset Description
This dataset contains two new benchmark corpora designed for low-resource languages spoken in the Democratic Republic of the Congo: The Lingala Read Speech Corpus LRSC, with 4.3 hours of labelled audio, and the Congolese Speech Radio Corpus CSRC, which offers 741 hours of unlabeled audio spanning four significant low-resource languages of the region (Lingala, Tshiluba… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/speech-recognition-congolese-languages.openslr-32-hq-SA-languages-Setswana
High quality TTS data for four South African languages - Setswana
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Setswana
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Setswana.mon_language_asr_audio
RFA Mon Language Voices
This dataset contains 14.8 hours of audio in the Mon language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Mon language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
This dataset was created by freococo.
The audio has been automatically segmented into 3,634 manageable chunks and… See the full description on the dataset page: https://huggingface.co/datasets/freococo/mon_language_asr_audio.rfa_shan_language_voices
RFA Shan Language Voices
This dataset contains 20.58 hours of audio in the Shan (Tai-Yai) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Shan language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
The audio has been automatically segmented into 5,047 manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/myandev/rfa_shan_language_voices.openslr-32-hq-SA-languages-Sesotho
High quality TTS data for four South African languages - Sesotho
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Sesotho
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/openslr-32-hq-SA-languages-Sesotho.rfa_rakhine_language_voices
RFA Rakhine Language Voices
This dataset contains 14.53 hours of audio in the Rakhine (Arakanese) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Rakhine language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
The audio has been automatically segmented into manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rfa_rakhine_language_voices.karenni_language_asr_audio
RFA Karenni (Kayah) Language Voices
This dataset contains 17 hours of audio in the Karenni (Kayah) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Karenni language family, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
This dataset was created by freococo.
The audio has been automatically segmented… See the full description on the dataset page: https://huggingface.co/datasets/freococo/karenni_language_asr_audio.samromur_syntheticSamrómur Synthetic consists of 72 hours of synthetized speech in Icelandic.whisper-eval-rare-languages-csv
Whisper 3 Large Evaluation on Mozilla Common Voice 17 Rare Languages (Enhanced Metrics)
Dataset Description
This enhanced dataset contains comprehensive evaluation results of OpenAI's Whisper 3 Large model on rare languages from Mozilla Common Voice 17, with extensive additional metrics for thorough ASR evaluation.
Key Features
Enhanced Error Metrics:
WER (Word Error Rate): Standard word-level error measurement
CER (Character Error Rate): Character-level error… See the full description on the dataset page: https://huggingface.co/datasets/norbertm/whisper-eval-rare-languages-csv.openslr-32-hq-SA-languages-Sesotho
High quality TTS data for four South African languages - Sesotho
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Sesotho
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Sesotho.rfa_shan_language_voices
RFA Shan Language Voices
This dataset contains 20.58 hours of audio in the Shan (Tai-Yai) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Shan language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
The audio has been automatically segmented into 5,047 manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rfa_shan_language_voices.openslr-32-hq-SA-languages-isiXhosa
High quality TTS data for four South African languages - isiXhosa
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - isiXhosa
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-isiXhosa.
