CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01oppThumbs /anv-ke-audioaudio100K<n<1M0 likes1.1k downloads12d agoHugging Face02dsfsi-anv /za-african-next-voicesgated Swivuriso: ZA-African Next Voices Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech, collected through ethical, community-centered processes. Dataset Paper: ArXiv - Work in Progress Language Coverage… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices.audioautomatic-speech-recognition100K<n<1M16 likes1k downloads7mo agoHugging Face03Anv-ke /Dholuogatedaudio100K<n<1M3 likes844 downloads7mo agoHugging Face04Anv-ke /kikuyugatedaudio100K<n<1M6 likes600 downloads7mo agoHugging Face05badrex /anv_data_ke_kikuyu_mergedaudio100K<n<1M0 likes409 downloads1y agoHugging Face06dsfsi-anv /multilingual-nchlt-dataset NCHLT Auxiliary Speech Corpus - Combined Multilingual Dataset Dataset Description This is a combined multilingual version of the NCHLT Auxiliary Speech Corpus, compiled by the Data Science for Social Impact (DSFSI) research group at the University of Pretoria to facilitate easier benchmarking and multi-language speech recognition research. The original auxiliary data was collected during the National Centre for Human Language Technology (NCHLT) project for the 11 official… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/multilingual-nchlt-dataset.audioautomatic-speech-recognition100K<n<1M1 likes379 downloads9mo agoHugging Face07badrex /anv_data_ke_kikuyu_scriptedaudio100K<n<1M0 likes363 downloads1y agoHugging Face08cryptpesa /anv-data-ke-somali-fullaudio10K<n<100K0 likes332 downloads5mo agoHugging Face09badrex /anv-data-ke-somali-fullaudio100K<n<1M1 likes326 downloads11mo agoHugging Face10MCAA1-MSU /anv_data_kegatedlanguage: ki so kln luo mas pretty_name: anv_ke ⚠️ IMPORTANT: Work in ProgressThis dataset is not final. Updates will continue through September 2025.Please use the latest version for attribution, benchmarking and publications. Overview African Next Voices: Pilot Data Collection in Kenya is part of a larger initiative to support African language speech technology. This project, funded by the Gates Foundation, is led by the KenCorpus Consortium, a coalition of Kenyan… See the full description on the dataset page: https://huggingface.co/datasets/MCAA1-MSU/anv_data_ke.audio100K<n<1M18 likes299 downloads7mo agoHugging Face11NjeriKahoro /anv-kikuyu-banking-subset-v2-part10audio1K<n<10K0 likes298 downloads23d agoHugging Face12Anv-ke /Kalenjingatedaudio10K<n<100K4 likes267 downloads7mo agoHugging Face13Anv-ke /Maasaigatedaudio10K<n<100K3 likes248 downloads7mo agoHugging Face14Anv-ke /Somaligatedaudio10K<n<100K5 likes198 downloads7mo agoHugging Face15NjeriKahoro /anv-kikuyu-banking-subset Anv-Kikuyu Banking Subset A domain-filtered subset of Kikuyu (Gĩkũyũ) speech data focused on banking and financial-transaction content, combined into a single repository with train, test, and validation splits. Source This dataset is a filtered subset of Anv-ke/kikuyu, part of the African Next Voices (ANV) collection. All audio, transcriptions, and underlying speaker data originate from that source dataset. Full credit for data collection belongs to the… See the full description on the dataset page: https://huggingface.co/datasets/NjeriKahoro/anv-kikuyu-banking-subset.audioautomatic-speech-recognition1K<n<10K1 likes105 downloads1mo agoHugging Face16badrex /anv-data-ke-somaliaudio10K<n<100K0 likes101 downloads11mo agoHugging Face17dsfsi-anv /za-african-next-voices-compressedgatedNote: This dataset is a compressed version of za-african-next-voices. It was compressed to .opus format using a 32k bitrate. Swivuriso: ZA-African Next Voices-Compressed Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices-compressed.audioautomatic-speech-recognition100K<n<1M1 likes86 downloads8mo agoHugging Face18badrex /anv-data-ke-somali-testaudion<1K0 likes27 downloads11mo agoHugging Face19badrex /anv-data-ke-kalenjin-evalaudio1K<n<10K0 likes19 downloads11mo agoHugging Face20dsfsi /anv_paper_sampleaudio1K<n<10K0 likes15 downloads1y agoHugging Face21MNG-4 /anv_scripted_multilingual-1audio1K<n<10K0 likes15 downloads3mo agoHugging Face22Kppwdfgu1 /anv-data-ke-somali-fullaudio10K<n<100K0 likes15 downloads3mo agoHugging Face23MNG-4 /anv_swahili_filteredaudio10K<n<100K0 likes9 downloads3mo agoHugging Face24evie-8 /anv-test-data-nt-ogggatedaudio10K<n<100K0 likes9 downloads3mo agoHugging Face25Kppwdfgu1 /anv_data_ke_kikuyu_scriptedaudio100K<n<1M0 likes8 downloads3mo agoHugging Face26cryptpesa /anv_data_ke_kikuyu_scriptedaudio100K<n<1M0 likes7 downloads5mo agoHugging Face27DigitalUmuganda /anv_test_data_nt_swahiligatedaudio10K<n<100K0 likes7 downloads4mo agoHugging Face28evie-8 /anv-test-data-ntgatedaudio10K<n<100K0 likes7 downloads3mo agoHugging Face29dsfsi /anv-za-sot-1h-sample-datasetgated Sesotho Sample Dataset - Next Voices-ZA (South Africa) - Multilingual Speech Dataset - Sesotho This dataset includes scripted and unscripted speech across various domains such as agriculture, health, finance, sports, transport, culture, society and general topics. It is primarily designed for automatic speech recognition (ASR). Use Restriction: The persons whose voices are included in this dataset, and the creators and owners of this dataset* do not give consent in… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi/anv-za-sot-1h-sample-dataset.audioautomatic-speech-recognitionn<1K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.