CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01softcatala /wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons. License identifiers are normalized to cc-zero, cc-by-4.0, cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self. This provides a richer alternative to Common Voice. Characteristics of the dataset: One or multiple speakers Different accents Different domain texts 761 audio files We found this dataset useful for audio tasks such as: Language detection Evaluation of STT systems New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.audioautomatic-speech-recognitionn<1K0 likes2.3k downloads2mo agoHugging Face02ymoslem /Wikimedia-Speech-Irish Dataset Details Synthetic audio dataset, created using Azure text-to-speech service. The bilingual text is a portion of the Wikimedia dataset, consisting of 7,545 text segments. The dataset includes two sets of audio data, one with a female voice (OrlaNeural) and the other with a male voice (ColmNeural). The speech data comprises approximately 34 hours and 23 minutes (34:23:12) spread across 15,090 utterances. Dataset Structure Dataset({ features: ['audio', 'text_ga'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Wikimedia-Speech-Irish.audioautomatic-speech-recognition10K<n<100K4 likes152 downloads2y agoHugging Face03Jaspernl /The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden Dataset Card for "The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden" Homepage: Spoken Wikipedia CorporaRepository: Bitbucket RepositoryPaper: Publication at LREC 2016Leaderboard: Interspeech 2018 PaperPoint of Contact: nats@nats.gitlab.io Dataset Summary The Spoken Wikipedia Corpora (SWC) is a collection of aligned spoken Wikipedia articles including articles in Dutch. It includes approximately 210 hours of audio, transcriptions, and metadata. The corpus is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Jaspernl/The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden.audioautomatic-speech-recognition10K<n<100K1 likes136 downloads2y agoHugging Face04ciempiess /wikipedia_spanish Dataset Card for wikipedia_spanish Dataset Summary According to the project page of the WikiProject Spoken Wikipedia: The WikiProject Spoken Wikipedia aims to produce recordings of Wikipedia articles being read aloud. Therefore, the WIKIPEDIA SPANISH CORPUS is a dataset created from the Spanish version of the WikiProject Spoken Wikipedia, called Wikipedia Grabada The WIKIPEDIA SPANISH CORPUS aims to be used in the Automatic Speech Recognition (ASR) task. It is a gender… See the full description on the dataset page: https://huggingface.co/datasets/ciempiess/wikipedia_spanish.audioautomatic-speech-recognition10K<n<100K1 likes83 downloads2y agoHugging Face05anaszil /Segmented-Moroccan-Darija-Wiki-Audio-Dataset Dataset Card for Segmented Moroccan Darija Wiki Dataset Dataset Summary This dataset provides short Moroccan Darija (Moroccan Arabic) speech segments derived from the atlasia/Moroccan-Darija-Wiki-Audio-Dataset.It is a cleaned and segmented version of the parent dataset, text-cleaned with Gemini 2.5-flash and processed using a fine-tuned Whisper model for Darija (to be open-sourced soon). Each audio is split into segments of up to 30 seconds to make it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/anaszil/Segmented-Moroccan-Darija-Wiki-Audio-Dataset.audioautomatic-speech-recognition1K<n<10K1 likes39 downloads1y agoHugging Face06pierluigic /WikIPA WikIPA Dataset Description WikIPA is a multilingual benchmark dataset designed for speech-to-IPA (STIPA) transcription, linking spoken audio with International Phonetic Alphabet (IPA) transcriptions. The dataset integrates two large-scale community-driven resources: WikiPron — human-curated IPA pronunciations extracted from Wiktionary Lingua Libre — crowdsourced recordings of spoken lexical items By connecting these two resources, WikIPA provides a dataset that links… See the full description on the dataset page: https://huggingface.co/datasets/pierluigic/WikIPA.audioautomatic-speech-recognition100K<n<1M2 likes31 downloads4mo agoHugging Face07BrunoHays /wikitongues-darija Wikitongues-Darija This is a small test dataset for Automatic Speech Recognition in Darija language, built from 2 captioned videos of the WikiTongues project: nawal anass Process: each webm video has been converted to monochannel 16khz wav files with ffmpeg : ffmpeg -i WIKITONGUES-_Nawal_speaking_Moroccan_Arabic.webm.1080p.vp9.webm -ar 16000 -ac 1 nawal.wav each audio has been cut in samples of less than 30 seconds audio according to the captions timestamps. The script may be… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/wikitongues-darija.audioautomatic-speech-recognitionn<1K1 likes22 downloads1y agoHugging Face08H20-sys /Segmented-Moroccan-Darija-Wiki-Audio-Dataset Dataset Card for Segmented Moroccan Darija Wiki Dataset Dataset Summary This dataset provides short Moroccan Darija (Moroccan Arabic) speech segments derived from the atlasia/Moroccan-Darija-Wiki-Audio-Dataset.It is a cleaned and segmented version of the parent dataset, text-cleaned with Gemini 2.5-flash and processed using a fine-tuned Whisper model for Darija (to be open-sourced soon). Each audio is split into segments of up to 30 seconds to make it suitable… See the full description on the dataset page: https://huggingface.co/datasets/H20-sys/Segmented-Moroccan-Darija-Wiki-Audio-Dataset.audioautomatic-speech-recognition1K<n<10K0 likes17 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.