datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ethiopian-languages-speech-dataset
Leyu Ethiopian Languages Speech Dataset
Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya.
Dataset Summary
Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti)
Total examples: 2750
License: CC-BY-4.0
Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/Benji-fish/ethiopian-languages-speech-dataset.mixed_multilingual_commonvoice_all_languages_100kBuild from mozilla commonvoice 13 using the script commited in this repo.
Used to teach a model to ignore languages that are not french
samromur_childrenThe Samrómur Children corpus contains more than 137000 validated speech-recordings uttered by Icelandic children.ePark_zu_yu_duan_wen_indigenous_language_essays
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_zu_yu_duan_wen_indigenous_language_essays
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_zu_yu_duan_wen_indigenous_language_essays.common_languageThe Common Language dataset without needing to run remote code, so it is compatible with datasets >= 4.0.0.
fleurs-r-neucodec-all-languages
FLEURS-R NeuCodec All Languages
FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102
locales, plus a speaker label FLEURS itself does not ship.
Layout
data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the
viewer shows).
audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members
named audio/{locale}/{split}/{id}.wav (the path column).
neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.mixed-language-detection-pilot-complete-sentences
Mixed-Language Speech Detection Pilot — Complete Sentences
This is the complete-sentence revision of a 6,000-clip binary
audio-classification pilot. label = 0 denotes one intended language and
label = 1 denotes more than one intended language. The covered languages are
Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani
(ckb), Arabic (ara), Persian (fas), and English (eng).
What changed
Earlier generation forced source transcripts into arbitrary… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-complete-sentences.mixed-language-detection-pilot-fleurs-voices
Mixed-Language Speech Detection Pilot — Native FLEURS Voices
This is the native-reference revision of a 6,000-clip binary
audio-classification pilot. label = 0 denotes one intended language and
label = 1 denotes more than one intended language. The covered languages are
Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani
(ckb), Arabic (ara), Persian (fas), and English (eng).
What changed in this revision
Synthetic speech is cloned from 36 real… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-fleurs-voices.mixed-language-detection-english-accented-vc
Mixed-Language Speech Detection Pilot
This dataset is a 6,000-clip binary audio-classification pilot for detecting
whether an utterance contains one language (label = 0) or more than one
language (label = 1). It covers Turkish (tur), Northern Kurdish/Kurmanji
(kmr), Central Kurdish/Sorani (ckb), Arabic (ara), Persian (fas), and
English (eng).
Dataset composition
Construction
Mixed
Monolingual
Total
Single-call OmniVoice
500
500
1,000
Segment-level… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-english-accented-vc.raddromur_asr
Dataset Card for raddromur_asr
Dataset Summary
The "Raddrómur Icelandic Speech 22.09" ("Raddrómur Corpus" for short) is an Icelandic corpus created by the Language and Voice Laboratory (LVL) at Reykjavík University (RU) in 2022. It is made out of radio podcasts mostly taken from RÚV (ruv.is).
Example Usage
The Raddrómur Corpus counts with the train split only. To load the training split pass its name as a config name:
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/language-and-voice-lab/raddromur_asr.ePark_qing_jing_zu_yu_contextual_indigenous_language
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_qing_jing_zu_yu_contextual_indigenous_language
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_qing_jing_zu_yu_contextual_indigenous_language.mixed_multilingual_commonvoice_all_languagesghanaian_languages_to_english_translation_and_transcription_datasetshan_language_asr_voices
⭐ A Voice for the Shan People: The SHAN Herald Agency Audio Archive
This is an extensive 306-hour audio dataset of the Shan (Tai-Yai) language, meticulously curated from the public broadcasts of the Shan Herald Agency for News (SHAN). For over two decades, SHAN has been a vital, independent voice for the people of Shan State, Myanmar, chronicling their stories of culture, politics, and the enduring struggle for federal democracy.
This archive stands as one of the largest publicly… See the full description on the dataset page: https://huggingface.co/datasets/freococo/shan_language_asr_voices.openslr-32-hq-SA-languages-Afrikaans
High quality TTS data for four South African languages - Afrikaans
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Afrikaans
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.10-languages-samplesThis dataset is suitable for evaluating language-identificatioan model
Afar-language-speech-transcription
Usage
This dataset is designed to support the development of text-to-speech (TTS) and speech-to-text (STT) systems for the Afar language. It can be integrated into web applications, mobile apps, desktop software, or other platforms that require natural-sounding Afar voice synthesis or accurate spoken language recognition.
For applications involving virtual avatars or voice personas, the following culturally appropriate voice names are recommended:
Female Voices: Emeli, Hanaawi… See the full description on the dataset page: https://huggingface.co/datasets/Charif-Ayfarah/Afar-language-speech-transcription.benji-ethiopian-languages-speech-dataset
Leyu Ethiopian Languages Speech Dataset
Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya.
Dataset Summary
Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti)
Total examples: 2750
License: CC-BY-4.0
Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/LeyuCompetition/benji-ethiopian-languages-speech-dataset.Low-Language-TTS
Low-Language-TTS
A held-out TTS/ASR test set for the long tail of
malaysia-ai/Multilingual-TTS:
50 of the lowest-resource languages in that corpus,
25 utterances each (1250 rows, 2.56 hours).
Every row carries the three things needed to score a NeuCodec speech-token model
without touching the parent corpus:
column
what
audio
the original clip, exactly as stored upstream (mostly mp3)
tokens
NeuCodec speech tokens at 50 tokens/s — the <|s_N|> ids the Multilingual-TTS… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Low-Language-TTS.openslr-32-hq-SA-languages
SLR32 – High Quality TTS Data for Four South African Languages
Identifier: SLR32License: CC BY-SA 4.0Source: https://www.openslr.org/32/
This dataset contains multi-speaker high quality transcribed audio data for four
languages of South Africa: Afrikaans (af_za), Sesotho (st_za),
Setswana (tn_za) and isiXhosa (xh_za).
The dataset consists of WAV files and a TSV file transcribing the audio.
In each folder the file line_index.tsv contains a FileID (which in turn
encodes the UserID)… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr-32-hq-SA-languages.speech-recognition-congolese-languages
Speech Recognition Datasets for Congolese Languages
Dataset Details
Dataset Description
This dataset contains two new benchmark corpora designed for low-resource languages spoken in the Democratic Republic of the Congo: The Lingala Read Speech Corpus LRSC, with 4.3 hours of labelled audio, and the Congolese Speech Radio Corpus CSRC, which offers 741 hours of unlabeled audio spanning four significant low-resource languages of the region (Lingala, Tshiluba… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/speech-recognition-congolese-languages.openslr-32-hq-SA-languages-Setswana
High quality TTS data for four South African languages - Setswana
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Setswana
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Setswana.rfa_shan_language_voices
RFA Shan Language Voices
This dataset contains 20.58 hours of audio in the Shan (Tai-Yai) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Shan language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
The audio has been automatically segmented into 5,047 manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/myandev/rfa_shan_language_voices.openslr-32-hq-SA-languages-Sesotho
High quality TTS data for four South African languages - Sesotho
Source - https://openslr.org/32/
Identifier: SLR32
Summary: Multi-speaker TTS data for four South African languages - Sesotho
License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
About this resource:
This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/openslr-32-hq-SA-languages-Sesotho.mon_language_asr_audio
RFA Mon Language Voices
This dataset contains 14.8 hours of audio in the Mon language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Mon language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
This dataset was created by freococo.
The audio has been automatically segmented into 3,634 manageable chunks and… See the full description on the dataset page: https://huggingface.co/datasets/freococo/mon_language_asr_audio.Language_Identificationmy-language-daily-speechsamromur_syntheticSamrómur Synthetic consists of 72 hours of synthetized speech in Icelandic.LanguageQA
Dataset Card for SAKURA-LanguageQA
This dataset contains the audio and the single/multi-hop questions/answers of the language track of the SAKURA benchmark from Interspeech 2025 paper, "SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information".
The fields of the dataset are:
file: The filename of the audio files.
audio: The audio recordings.
attribute_label: The attribute labels (i.e., the language spoken in the speech) of the audio… See the full description on the dataset page: https://huggingface.co/datasets/SLLM-multi-hop/LanguageQA.rfa_rakhine_language_voices
RFA Rakhine Language Voices
This dataset contains 14.53 hours of audio in the Rakhine (Arakanese) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Rakhine language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
The audio has been automatically segmented into manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rfa_rakhine_language_voices.
