datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
goan-konkani-speech
Goan Konkani Speech (Romi)
47,365 audio clips, 108.5 hours of Goan Konkani (ISO 639-3 gom)
speech from Goan television news, transcribed in Romi Konkani - Konkani written in
the Roman script.
Konkani is a low-resource language with very little public speech data. This is
assembled from broadcast news, so it is real spoken Konkani: studio anchors, field
reporters, phone interviews, and the Konkani-English code-switching that Goan speakers
actually use.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/goan-konkani-speech.Goal-Dataset_en_innew-moore-speech-clean
Moore Speech Proverbs: A Parallel Audio-Text Corpus for Mooré and French
The Moore Speech Proverbs dataset is a bilingual audio-text corpus of traditional proverbs in Mooré and French, designed for research and academic purposes in low-resource speech and language processing.
It is intended primarily for academic or research purposes in text-to-speech (TTS) and automatic speech recognition (ASR) for Mooré language.
[!NOTE]
⚠️ Access is gated. To request access, please read the… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/new-moore-speech-clean.new-dioula-speech-clean
Dioula Speech Corpus: A Parallel Audio-Text Dataset for Dioula and French
The Dioula Speech Corpus is a bilingual audio-text corpus designed for research and academic purposes in low-resource speech and language processing. It is intended primarily to support the development of Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) models for the Dioula language.
⚠️ Access is gated. To request access, please read the policy below.🛑 TLDR: For safety and traceability reasons… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/new-dioula-speech-clean.vaani-goa_northsouthgoa-cleanedgoai-moore-speech-contes
Moore Speech Contes: A Spoken Corpus of Traditional Mooré Stories
The Moore Speech Contes dataset is a collection of spoken folk stories (contes) in Mooré, designed for research and academic purposes in low-resource speech and language processing.
It is intended primarily for academic or research purposes in text-to-speech (TTS) and automatic speech recognition (ASR) for Mooré language (ISO 639-3: mos).
[!NOTE]
⚠️ Access is gated. To request access, please read the policy below.🚩… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/goai-moore-speech-contes.GOAI-MooreSpeechCorpora
Moore Speech Corpora: A Cleaned, Denoised Audio-Text Dataset for Mooré TTS and ASR
The Moore Speech Corpora is a collection of aligned audio and text in Mooré, gathered from publicly available sources. This unified corpus is curated for research and academic purposes in low-resource speech and language processing, especially for text-to-speech (TTS) and automatic speech recognition (ASR) in the Mooré language (ISO 639-3: mos).
Mooré is under-represented in current speech corpora and… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/GOAI-MooreSpeechCorpora.goai-dioula-speechmoore-speech-bible
Moore Speech Bible: A Curated Audio-Text Dataset for Mooré TTS and ASR
The Moore Speech Bible dataset is a collection of aligned audio and text in Mooré, gathered from publicly available religious sources. This corpus is curated for research and academic purposes in low-resource speech and language processing, especially for text-to-speech (TTS) and automatic speech recognition (ASR) in the Mooré language (ISO 639-3: mos).
Mooré remains under-represented in current speech… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/moore-speech-bible.goai-moore-speech-devinettes
Moore Speech Devinettes: A Spoken Riddle Dataset in Mooré
The Moore Speech Devinettes dataset is a spoken collection of traditional Mooré riddles, created for academic and research use in low-resource speech and language technologies.
It is designed to support work in text-to-speech (TTS), automatic speech recognition (ASR), and oral tradition modeling for the Mooré language (ISO 639-3: mos).
[!NOTE]
⚠️ Access is gated. To request access, please read the policy below.
🚩 TLDR: For… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/goai-moore-speech-devinettes.
