CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Reubencf /goan-konkani-speech Goan Konkani Speech (Romi) 47,365 audio clips, 108.5 hours of Goan Konkani (ISO 639-3 gom) speech from Goan television news, transcribed in Romi Konkani - Konkani written in the Roman script. Konkani is a low-resource language with very little public speech data. This is assembled from broadcast news, so it is real spoken Konkani: studio anchors, field reporters, phone interviews, and the Konkani-English code-switching that Goan speakers actually use. Contents… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/goan-konkani-speech.audioautomatic-speech-recognition10K<n<100K0 likes178 downloads8d agoHugging Face02VoiceArena /Goal-Dataset_en_inaudion<1K2 likes174 downloads9d agoHugging Face03goaicorp /new-moore-speech-cleangated Moore Speech Proverbs: A Parallel Audio-Text Corpus for Mooré and French The Moore Speech Proverbs dataset is a bilingual audio-text corpus of traditional proverbs in Mooré and French, designed for research and academic purposes in low-resource speech and language processing. It is intended primarily for academic or research purposes in text-to-speech (TTS) and automatic speech recognition (ASR) for Mooré language. [!NOTE] ⚠️ Access is gated. To request access, please read the… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/new-moore-speech-clean.audiotext-to-speech1K<n<10K1 likes18 downloads5mo agoHugging Face04goaicorp /new-dioula-speech-cleangated Dioula Speech Corpus: A Parallel Audio-Text Dataset for Dioula and French The Dioula Speech Corpus is a bilingual audio-text corpus designed for research and academic purposes in low-resource speech and language processing. It is intended primarily to support the development of Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) models for the Dioula language. ⚠️ Access is gated. To request access, please read the policy below.🛑 TLDR: For safety and traceability reasons… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/new-dioula-speech-clean.audiotext-to-speech10K<n<100K1 likes15 downloads5mo agoHugging Face05abar-uwc /vaani-goa_northsouthgoa-cleanedaudio1K<n<10K0 likes14 downloads1y agoHugging Face06goaicorp /goai-moore-speech-contesgated Moore Speech Contes: A Spoken Corpus of Traditional Mooré Stories The Moore Speech Contes dataset is a collection of spoken folk stories (contes) in Mooré, designed for research and academic purposes in low-resource speech and language processing. It is intended primarily for academic or research purposes in text-to-speech (TTS) and automatic speech recognition (ASR) for Mooré language (ISO 639-3: mos). [!NOTE] ⚠️ Access is gated. To request access, please read the policy below.🚩… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/goai-moore-speech-contes.audiotext-to-speech1K<n<10K1 likes12 downloads6mo agoHugging Face07goaicorp /GOAI-MooreSpeechCorporagated Moore Speech Corpora: A Cleaned, Denoised Audio-Text Dataset for Mooré TTS and ASR The Moore Speech Corpora is a collection of aligned audio and text in Mooré, gathered from publicly available sources. This unified corpus is curated for research and academic purposes in low-resource speech and language processing, especially for text-to-speech (TTS) and automatic speech recognition (ASR) in the Mooré language (ISO 639-3: mos). Mooré is under-represented in current speech corpora and… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/GOAI-MooreSpeechCorpora.audiotext-to-speech10K<n<100K3 likes9 downloads6mo agoHugging Face08goaicorp /goai-dioula-speechgatedaudio10K<n<100K1 likes8 downloads1y agoHugging Face09goaicorp /moore-speech-biblegated Moore Speech Bible: A Curated Audio-Text Dataset for Mooré TTS and ASR The Moore Speech Bible dataset is a collection of aligned audio and text in Mooré, gathered from publicly available religious sources. This corpus is curated for research and academic purposes in low-resource speech and language processing, especially for text-to-speech (TTS) and automatic speech recognition (ASR) in the Mooré language (ISO 639-3: mos). Mooré remains under-represented in current speech… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/moore-speech-bible.audio10K<n<100K1 likes5 downloads6mo agoHugging Face10goaicorp /goai-moore-speech-devinettesgated Moore Speech Devinettes: A Spoken Riddle Dataset in Mooré The Moore Speech Devinettes dataset is a spoken collection of traditional Mooré riddles, created for academic and research use in low-resource speech and language technologies. It is designed to support work in text-to-speech (TTS), automatic speech recognition (ASR), and oral tradition modeling for the Mooré language (ISO 639-3: mos). [!NOTE] ⚠️ Access is gated. To request access, please read the policy below. 🚩 TLDR: For… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/goai-moore-speech-devinettes.audiotext-to-speechn<1K1 likes3 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.