CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Benji-fish /ethiopian-languages-speech-dataset Leyu Ethiopian Languages Speech Dataset Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya. Dataset Summary Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti) Total examples: 2750 License: CC-BY-4.0 Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/Benji-fish/ethiopian-languages-speech-dataset.audioautomatic-speech-recognition1K<n<10K0 likes982 downloads1mo agoHugging Face02language-and-voice-lab /samromur_childrenThe Samrómur Children corpus contains more than 137000 validated speech-recordings uttered by Icelandic children.audioautomatic-speech-recognition10K<n<100K9 likes570 downloads3y agoHugging Face03FormosanBank /ePark_zu_yu_duan_wen_indigenous_language_essays FormosanBank publication status This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card. FormosanBank/ePark_zu_yu_duan_wen_indigenous_language_essays Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_zu_yu_duan_wen_indigenous_language_essays.audioautomatic-speech-recognition1K<n<10K0 likes436 downloads2mo agoHugging Face04malaysia-ai /fleurs-r-neucodec-all-languages FLEURS-R NeuCodec All Languages FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102 locales, plus a speaker label FLEURS itself does not ship. Layout data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the viewer shows). audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members named audio/{locale}/{split}/{id}.wav (the path column). neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.audiotext-to-speech100K<n<1M4 likes245 downloads10d agoHugging Face05language-and-voice-lab /raddromur_asr Dataset Card for raddromur_asr Dataset Summary The "Raddrómur Icelandic Speech 22.09" ("Raddrómur Corpus" for short) is an Icelandic corpus created by the Language and Voice Laboratory (LVL) at Reykjavík University (RU) in 2022. It is made out of radio podcasts mostly taken from RÚV (ruv.is). Example Usage The Raddrómur Corpus counts with the train split only. To load the training split pass its name as a config name: from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/language-and-voice-lab/raddromur_asr.audioautomatic-speech-recognition10K<n<100K3 likes180 downloads2y agoHugging Face06FormosanBank /ePark_qing_jing_zu_yu_contextual_indigenous_language FormosanBank publication status This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card. FormosanBank/ePark_qing_jing_zu_yu_contextual_indigenous_language Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_qing_jing_zu_yu_contextual_indigenous_language.audioautomatic-speech-recognition10K<n<100K0 likes134 downloads2mo agoHugging Face07freococo /shan_language_asr_voices ⭐ A Voice for the Shan People: The SHAN Herald Agency Audio Archive This is an extensive 306-hour audio dataset of the Shan (Tai-Yai) language, meticulously curated from the public broadcasts of the Shan Herald Agency for News (SHAN). For over two decades, SHAN has been a vital, independent voice for the people of Shan State, Myanmar, chronicling their stories of culture, politics, and the enduring struggle for federal democracy. This archive stands as one of the largest publicly… See the full description on the dataset page: https://huggingface.co/datasets/freococo/shan_language_asr_voices.audioautomatic-speech-recognition10K<n<100K2 likes92 downloads1y agoHugging Face08voice-biomarkers /openslr-32-hq-SA-languages-Afrikaans High quality TTS data for four South African languages - Afrikaans Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Afrikaans License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.audioautomatic-speech-recognition1K<n<10K5 likes90 downloads2y agoHugging Face09LeyuCompetition /benji-ethiopian-languages-speech-dataset Leyu Ethiopian Languages Speech Dataset Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya. Dataset Summary Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti) Total examples: 2750 License: CC-BY-4.0 Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/LeyuCompetition/benji-ethiopian-languages-speech-dataset.audioautomatic-speech-recognition1K<n<10K0 likes71 downloads21d agoHugging Face10deepdml /openslr-32-hq-SA-languages SLR32 – High Quality TTS Data for Four South African Languages Identifier: SLR32License: CC BY-SA 4.0Source: https://www.openslr.org/32/ This dataset contains multi-speaker high quality transcribed audio data for four languages of South Africa: Afrikaans (af_za), Sesotho (st_za), Setswana (tn_za) and isiXhosa (xh_za). The dataset consists of WAV files and a TSV file transcribing the audio. In each folder the file line_index.tsv contains a FileID (which in turn encodes the UserID)… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr-32-hq-SA-languages.audioautomatic-speech-recognition1K<n<10K0 likes52 downloads7mo agoHugging Face11Svngoku /speech-recognition-congolese-languages Speech Recognition Datasets for Congolese Languages Dataset Details Dataset Description This dataset contains two new benchmark corpora designed for low-resource languages spoken in the Democratic Republic of the Congo: The Lingala Read Speech Corpus LRSC, with 4.3 hours of labelled audio, and the Congolese Speech Radio Corpus CSRC, which offers 741 hours of unlabeled audio spanning four significant low-resource languages of the region (Lingala, Tshiluba… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/speech-recognition-congolese-languages.audioautomatic-speech-recognition1K<n<10K4 likes45 downloads2y agoHugging Face12voice-biomarkers /openslr-32-hq-SA-languages-Setswana High quality TTS data for four South African languages - Setswana Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Setswana License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Setswana.audioautomatic-speech-recognition1K<n<10K1 likes44 downloads2y agoHugging Face13freococo /mon_language_asr_audio RFA Mon Language Voices This dataset contains 14.8 hours of audio in the Mon language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Mon language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks. This dataset was created by freococo. The audio has been automatically segmented into 3,634 manageable chunks and… See the full description on the dataset page: https://huggingface.co/datasets/freococo/mon_language_asr_audio.audioautomatic-speech-recognition1K<n<10K0 likes43 downloads1y agoHugging Face14freococo /rfa_rakhine_language_voices RFA Rakhine Language Voices This dataset contains 14.53 hours of audio in the Rakhine (Arakanese) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Rakhine language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks. The audio has been automatically segmented into manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rfa_rakhine_language_voices.audioautomatic-speech-recognition1K<n<10K0 likes39 downloads1y agoHugging Face15myandev /rfa_shan_language_voices RFA Shan Language Voices This dataset contains 20.58 hours of audio in the Shan (Tai-Yai) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Shan language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks. The audio has been automatically segmented into 5,047 manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/myandev/rfa_shan_language_voices.audioautomatic-speech-recognition0 likes39 downloads5d agoHugging Face16freococo /karenni_language_asr_audio RFA Karenni (Kayah) Language Voices This dataset contains 17 hours of audio in the Karenni (Kayah) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Karenni language family, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks. This dataset was created by freococo. The audio has been automatically segmented… See the full description on the dataset page: https://huggingface.co/datasets/freococo/karenni_language_asr_audio.audioautomatic-speech-recognition1K<n<10K0 likes37 downloads1y agoHugging Face17Max5ive /openslr-32-hq-SA-languages-Sesotho High quality TTS data for four South African languages - Sesotho Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Sesotho License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/openslr-32-hq-SA-languages-Sesotho.audioautomatic-speech-recognition1K<n<10K0 likes36 downloads6mo agoHugging Face18freococo /rfa_shan_language_voices RFA Shan Language Voices This dataset contains 20.58 hours of audio in the Shan (Tai-Yai) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Shan language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks. The audio has been automatically segmented into 5,047 manageable chunks and prepared in the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rfa_shan_language_voices.audioautomatic-speech-recognition1K<n<10K0 likes30 downloads1y agoHugging Face19voice-biomarkers /openslr-32-hq-SA-languages-Sesotho High quality TTS data for four South African languages - Sesotho Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Sesotho License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Sesotho.audioautomatic-speech-recognition1K<n<10K2 likes29 downloads2y agoHugging Face20language-and-voice-lab /samromur_syntheticSamrómur Synthetic consists of 72 hours of synthetized speech in Icelandic.audioautomatic-speech-recognition10K<n<100K1 likes28 downloads2y agoHugging Face21voice-biomarkers /openslr-32-hq-SA-languages-isiXhosa High quality TTS data for four South African languages - isiXhosa Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - isiXhosa License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-isiXhosa.audioautomatic-speech-recognition1K<n<10K1 likes20 downloads2y agoHugging Face22African-Languages-Lab /kasagadigated Kasagadi — Ghanaian Radio Broadcast Fact-Check Dataset This is a multilingual dataset of transcribed, translated, and AI fact-checked segments from live radio broadcasts across Ghana. It is a Ghanaian initiative, covering Twi-language broadcasts from two Ghanaian FM stations and Hausa-language broadcasts from a third Ghanaian FM station serving Ghana's Zongo communities. Dataset Summary Station Language Broadcasts Segments Hours Date Range Angel FM Twi… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/kasagadi.audiotext-classification100K<n<1M0 likes17 downloads3mo agoHugging Face23African-Languages-Lab /all-lab-speechgated All Lab Speech Cleaned African-language speech with embedded, playable audio (HF Audio). One config per language (<lang> = transcribed, <lang>_manifest = audio-only); the audio column sits right after audio_id and plays in the dataset viewer. Splits (train/validation/test) come from the source split labels. from datasets import load_dataset ds = load_dataset("African-Languages-Lab/all-lab-speech", "afrikaans") Columns audio_id, audio (playable), transcript… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/all-lab-speech.audioautomatic-speech-recognition10M<n<100M2 likes4 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.