CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Benji-fish /ethiopian-languages-speech-dataset Leyu Ethiopian Languages Speech Dataset Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya. Dataset Summary Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti) Total examples: 2750 License: CC-BY-4.0 Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/Benji-fish/ethiopian-languages-speech-dataset.audioautomatic-speech-recognition1K<n<10K0 likes1k downloads1mo agoHugging Face02malaysia-ai /fleurs-r-neucodec-all-languages FLEURS-R NeuCodec All Languages FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102 locales, plus a speaker label FLEURS itself does not ship. Layout data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the viewer shows). audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members named audio/{locale}/{split}/{id}.wav (the path column). neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.audiotext-to-speech100K<n<1M4 likes256 downloads9d agoHugging Face03voice-biomarkers /openslr-32-hq-SA-languages-Afrikaans High quality TTS data for four South African languages - Afrikaans Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Afrikaans License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Afrikaans.audioautomatic-speech-recognition1K<n<10K5 likes89 downloads2y agoHugging Face04LeyuCompetition /benji-ethiopian-languages-speech-dataset Leyu Ethiopian Languages Speech Dataset Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya. Dataset Summary Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti) Total examples: 2750 License: CC-BY-4.0 Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/LeyuCompetition/benji-ethiopian-languages-speech-dataset.audioautomatic-speech-recognition1K<n<10K0 likes69 downloads21d agoHugging Face05deepdml /openslr-32-hq-SA-languages SLR32 – High Quality TTS Data for Four South African Languages Identifier: SLR32License: CC BY-SA 4.0Source: https://www.openslr.org/32/ This dataset contains multi-speaker high quality transcribed audio data for four languages of South Africa: Afrikaans (af_za), Sesotho (st_za), Setswana (tn_za) and isiXhosa (xh_za). The dataset consists of WAV files and a TSV file transcribing the audio. In each folder the file line_index.tsv contains a FileID (which in turn encodes the UserID)… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/openslr-32-hq-SA-languages.audioautomatic-speech-recognition1K<n<10K0 likes50 downloads7mo agoHugging Face06Svngoku /speech-recognition-congolese-languages Speech Recognition Datasets for Congolese Languages Dataset Details Dataset Description This dataset contains two new benchmark corpora designed for low-resource languages spoken in the Democratic Republic of the Congo: The Lingala Read Speech Corpus LRSC, with 4.3 hours of labelled audio, and the Congolese Speech Radio Corpus CSRC, which offers 741 hours of unlabeled audio spanning four significant low-resource languages of the region (Lingala, Tshiluba… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/speech-recognition-congolese-languages.audioautomatic-speech-recognition1K<n<10K4 likes47 downloads2y agoHugging Face07voice-biomarkers /openslr-32-hq-SA-languages-Setswana High quality TTS data for four South African languages - Setswana Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Setswana License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Setswana.audioautomatic-speech-recognition1K<n<10K1 likes43 downloads2y agoHugging Face08Max5ive /openslr-32-hq-SA-languages-Sesotho High quality TTS data for four South African languages - Sesotho Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Sesotho License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/openslr-32-hq-SA-languages-Sesotho.audioautomatic-speech-recognition1K<n<10K0 likes36 downloads6mo agoHugging Face09norbertm /whisper-eval-rare-languages-csv Whisper 3 Large Evaluation on Mozilla Common Voice 17 Rare Languages (Enhanced Metrics) Dataset Description This enhanced dataset contains comprehensive evaluation results of OpenAI's Whisper 3 Large model on rare languages from Mozilla Common Voice 17, with extensive additional metrics for thorough ASR evaluation. Key Features Enhanced Error Metrics: WER (Word Error Rate): Standard word-level error measurement CER (Character Error Rate): Character-level error… See the full description on the dataset page: https://huggingface.co/datasets/norbertm/whisper-eval-rare-languages-csv.automatic-speech-recognition0 likes33 downloads1y agoHugging Face10voice-biomarkers /openslr-32-hq-SA-languages-Sesotho High quality TTS data for four South African languages - Sesotho Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - Sesotho License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio. In… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-Sesotho.audioautomatic-speech-recognition1K<n<10K2 likes26 downloads2y agoHugging Face11voice-biomarkers /openslr-32-hq-SA-languages-isiXhosa High quality TTS data for four South African languages - isiXhosa Source - https://openslr.org/32/ Identifier: SLR32 Summary: Multi-speaker TTS data for four South African languages - isiXhosa License: Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) About this resource: This data set contains multi-speaker high quality transcribed audio data for four languages of South Africa. The data set consists of wave files, and a TSV file transcribing the audio.… See the full description on the dataset page: https://huggingface.co/datasets/voice-biomarkers/openslr-32-hq-SA-languages-isiXhosa.audioautomatic-speech-recognition1K<n<10K1 likes19 downloads2y agoHugging Face12FrancophonIA /17-minute-world-languages_allemande [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/allemande/ Site à scrapper translation0 likes16 downloads1y agoHugging Face13FrancophonIA /17-minute-world-languages_georgien [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/géorgienne/ Site à scrapper translation0 likes16 downloads1y agoHugging Face14FrancophonIA /17-minute-world-languages_malais [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/malaise/ Site à scrapper translation0 likes16 downloads1y agoHugging Face15FrancophonIA /17-minute-world-languages_jordanien [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/jordanienne/ Site à scrapper translation0 likes15 downloads1y agoHugging Face16FrancophonIA /17-minute-world-languages_malgache [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/malgache/ Site à scrapper translation0 likes15 downloads1y agoHugging Face17FrancophonIA /17-minute-world-languages_persan [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/persane/ Site à scrapper translation0 likes15 downloads1y agoHugging Face18African-Languages-Lab /kasagadigated Kasagadi — Ghanaian Radio Broadcast Fact-Check Dataset This is a multilingual dataset of transcribed, translated, and AI fact-checked segments from live radio broadcasts across Ghana. It is a Ghanaian initiative, covering Twi-language broadcasts from two Ghanaian FM stations and Hausa-language broadcasts from a third Ghanaian FM station serving Ghana's Zongo communities. Dataset Summary Station Language Broadcasts Segments Hours Date Range Angel FM Twi… See the full description on the dataset page: https://huggingface.co/datasets/African-Languages-Lab/kasagadi.audiotext-classification100K<n<1M0 likes13 downloads3mo agoHugging Face19FrancophonIA /17-minute-world-languages_bengali [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/bengali/ Site à scrapper translation0 likes12 downloads1y agoHugging Face20FrancophonIA /17-minute-world-languages_mongol [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/mongole/ Site à scrapper translation0 likes12 downloads1y agoHugging Face21FrancophonIA /17-minute-world-languages_amharique [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/amharique/ Site à scrapper translation0 likes11 downloads1y agoHugging Face22FrancophonIA /17-minute-world-languages_afrikaans [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/afrikaans/ Site à scrapper translation0 likes9 downloads1y agoHugging Face23FrancophonIA /17-minute-world-languages_albanaise [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/albanaise/ Site à scrapper translation0 likes9 downloads1y agoHugging Face24FrancophonIA /17-minute-world-languages_azeri [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/azéri/ Site à scrapper translation0 likes9 downloads1y agoHugging Face25FrancophonIA /17-minute-world-languages_bosniaque [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/bosniaque/ Site à scrapper translation0 likes9 downloads1y agoHugging Face26FrancophonIA /17-minute-world-languages_dari [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/dari/ Site à scrapper translation0 likes9 downloads1y agoHugging Face27FrancophonIA /17-minute-world-languages_egyptien [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/égyptienne/ Site à scrapper translation0 likes9 downloads1y agoHugging Face28FrancophonIA /17-minute-world-languages_lingala [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/lingala/ Site à scrapper translation0 likes9 downloads1y agoHugging Face29FrancophonIA /17-minute-world-languages_slovene [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/slovène/ Site à scrapper translation0 likes9 downloads1y agoHugging Face30FrancophonIA /17-minute-world-languages_swahili [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/swahilie/ Site à scrapper translation0 likes9 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.