CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01badayvedat /VCTKaudio10K<n<100K6 likes1k downloads2y agoHugging Face02badrex /malagasy-speech-fullaudio10K<n<100K2 likes447 downloads11mo agoHugging Face03badrex /anv_data_ke_kikuyu_mergedaudio100K<n<1M0 likes336 downloads1y agoHugging Face04badrex /anv_data_ke_kikuyu_scriptedaudio100K<n<1M0 likes314 downloads1y agoHugging Face05badrex /kinyarwanda-speech-500h Kinyarwanda Automatic Speech Recognition Dataset Dataset Description This dataset contains 500 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition. Dataset Details Language: Kinyarwanda (rw) Task: Automatic Speech Recognition Size: ~500 hours of transcribed speech Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-500h.audioautomatic-speech-recognition10K<n<100K0 likes300 downloads1y agoHugging Face06badrex /anv-data-ke-somali-fullaudio100K<n<1M1 likes289 downloads11mo agoHugging Face07badrex /ethiopian-speech-flat Ethio Speech Copus — Afrivoices Ethiopian 📌 Overview The Ethio Speech Corpus dataset is a multilingual speech corpus containing audio–text pairs across five Ethiopian languages. It is designed to support the development of speech-to-text technologies for low-resource languages. This dataset is part of the Afrivoices initiative — a collaborative effort to create a large-scale ASR dataset for African languages. The broader goal of the initiative is to collect 600 hours… See the full description on the dataset page: https://huggingface.co/datasets/badrex/ethiopian-speech-flat.audioautomatic-speech-recognition100K<n<1M2 likes279 downloads8mo agoHugging Face08badrex /kinyarwanda-speech-1000h Kinyarwanda Automatic Speech Recognition Dataset Dataset Description This dataset contains ~1000 hours of transcribed Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track B competition. Dataset Details Language: Kinyarwanda (rw) Task: Automatic Speech Recognition Size: ~1000 hours of transcribed speech Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-1000h.audioautomatic-speech-recognition100K<n<1M0 likes211 downloads1y agoHugging Face09badrex /swahili-speech-400hraudioautomatic-speech-recognition10K<n<100K1 likes206 downloads11mo agoHugging Face10badrex /shona-speechaudio10K<n<100K8 likes189 downloads11mo agoHugging Face11badayvedat /LJSpeech-1.1 The LJ Speech Dataset Version 1.1 July 5, 2017 https://keithito.com/LJ-Speech-Dataset OVERVIEW This is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading passages from 7 non-fiction books. A transcription is provided for each clip. Clips vary in length from 1 to 10 seconds and have a total length of approximately 24 hours. The texts were published between 1884 and 1964, and are in the public domain. The audio was recorded in… See the full description on the dataset page: https://huggingface.co/datasets/badayvedat/LJSpeech-1.1.audio10K<n<100K0 likes168 downloads2y agoHugging Face12badrex /wolaytta-speechaudio10K<n<100K0 likes116 downloads11mo agoHugging Face13BadiniSpeechNLP /fleurs-badini FLEURS-Badini Dataset Summary FLEURS-Badini is a speech dataset for the Badini dialect of Northern Kurdish, designed for research in: Automatic Speech Recognition (ASR) Speech-to-Text Translation (S2TT) It is a dialect-specific extension of the FLEURS benchmark, providing aligned speech–text–translation data for a low-resource language variant. The dataset contains 5,224 utterances (~15h40m) recorded from 45 speakers. Supported Tasks Automatic Speech… See the full description on the dataset page: https://huggingface.co/datasets/BadiniSpeechNLP/fleurs-badini.audio1K<n<10K1 likes112 downloads5mo agoHugging Face14BadiniSpeechNLP /badini-asr-benchmark Dataset Card for "badini-asr-benchmark" More Information needed audio1K<n<10K1 likes108 downloads5mo agoHugging Face15badrex /amharic-speechaudio10K<n<100K0 likes107 downloads11mo agoHugging Face16badrex /MADIS5-spoken-arabic-dialects Dataset Overview MADIS-5 (Multi-domain Arabic Dialect Identification in Speech) is a manually curated dataset designed to facilitate evaluation of cross-domain robustness of Arabic Dialect Identification (ADI) systems. This dataset provides a comprehensive benchmark for testing out-of-domain generalization across different speech domains with diverse recording conditions and speaking styles. Dataset Statistics Total Duration: ~12 hours of speech Total… See the full description on the dataset page: https://huggingface.co/datasets/badrex/MADIS5-spoken-arabic-dialects.audioaudio-classification1K<n<10K0 likes106 downloads1y agoHugging Face17badrex /kalenjin-speech-fullaudio10K<n<100K0 likes101 downloads11mo agoHugging Face18badrex /anv-data-ke-somaliaudio10K<n<100K0 likes86 downloads11mo agoHugging Face19badrex /arabic-speech-SADA22-MSA Dataset Card for SADA (Saudi Audio Dataset for Arabic) ⚠️ Caution This is only the portion of the SADA dataset where the speaker dialect is Modern Standard Arabic (MSA). To access full dataset, you should check this link. Dataset Summary The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-MSA.audioautomatic-speech-recognition1K<n<10K2 likes83 downloads1y agoHugging Face20badrex /afaan_oromo-speechaudio10K<n<100K0 likes83 downloads11mo agoHugging Face21badrex /fulani-speechaudio10K<n<100K0 likes68 downloads1y agoHugging Face22badrex /ethiopian-speechaudio10K<n<100K1 likes66 downloads11mo agoHugging Face23computeram /badini-tts-samplesaudion<1K0 likes57 downloads6d agoHugging Face24badrex /tigrinya-speechaudio10K<n<100K0 likes53 downloads11mo agoHugging Face25badrex /arabic-speech-SADA22-Khaliji Dataset Card for SADA (Saudi Audio Dataset for Arabic) ⚠️ Caution This is only the portion of the SADA dataset where the speaker dialect is Khaliji. To access full dataset, you should check this link. Dataset Summary The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of… See the full description on the dataset page: https://huggingface.co/datasets/badrex/arabic-speech-SADA22-Khaliji.audioautomatic-speech-recognition10K<n<100K5 likes49 downloads1y agoHugging Face26badrex /waxalNLP-ethiopic-finalaudio100K<n<1M0 likes38 downloads7mo agoHugging Face27LennyBijan /BA_Datensatz_V2audioautomatic-speech-recognition1K<n<10K1 likes37 downloads3y agoHugging Face28Hindy /Bad_Therapeutic_Musicaudio100K<n<1M0 likes34 downloads2y agoHugging Face29badrex /anv-data-ke-somali-testaudion<1K0 likes30 downloads11mo agoHugging Face30badrex /sidama-speechaudio10K<n<100K0 likes29 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.