datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
goan-konkani-speech
Goan Konkani Speech (Romi)
47,365 audio clips, 108.5 hours of Goan Konkani (ISO 639-3 gom)
speech from Goan television news, transcribed in Romi Konkani - Konkani written in
the Roman script.
Konkani is a low-resource language with very little public speech data. This is
assembled from broadcast news, so it is real spoken Konkani: studio anchors, field
reporters, phone interviews, and the Konkani-English code-switching that Goan speakers
actually use.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/goan-konkani-speech.new-moore-speech-clean
Moore Speech Proverbs: A Parallel Audio-Text Corpus for Mooré and French
The Moore Speech Proverbs dataset is a bilingual audio-text corpus of traditional proverbs in Mooré and French, designed for research and academic purposes in low-resource speech and language processing.
It is intended primarily for academic or research purposes in text-to-speech (TTS) and automatic speech recognition (ASR) for Mooré language.
[!NOTE]
⚠️ Access is gated. To request access, please read the… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/new-moore-speech-clean.new-dioula-speech-clean
Dioula Speech Corpus: A Parallel Audio-Text Dataset for Dioula and French
The Dioula Speech Corpus is a bilingual audio-text corpus designed for research and academic purposes in low-resource speech and language processing. It is intended primarily to support the development of Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) models for the Dioula language.
⚠️ Access is gated. To request access, please read the policy below.🛑 TLDR: For safety and traceability reasons… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/new-dioula-speech-clean.goan-konkani-raw-audio
Goan Konkani Raw Audio
Private raw-ingestion index. Audio objects are preserved in their original YouTube audio container in the private Reubencf/goan-konkani-raw-audio Storage Bucket. Cleaning, repeated-Mass detection, song removal, segmentation, and transcription are derived stages; the raw source objects are not overwritten. Access and reuse remain subject to the source owners’ rights and permissions.
data/pilot_metadata.jsonl records checksums, durations, source URLs, and… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/goan-konkani-raw-audio.goai-moore-speech-contes
Moore Speech Contes: A Spoken Corpus of Traditional Mooré Stories
The Moore Speech Contes dataset is a collection of spoken folk stories (contes) in Mooré, designed for research and academic purposes in low-resource speech and language processing.
It is intended primarily for academic or research purposes in text-to-speech (TTS) and automatic speech recognition (ASR) for Mooré language (ISO 639-3: mos).
[!NOTE]
⚠️ Access is gated. To request access, please read the policy below.🚩… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/goai-moore-speech-contes.goatis-transcripts
Goatis / Sv3rige Video Transcripts
Full transcripts of 1,383 videos (~487 hours, ~3.9M words) from the
YouTube channels of Goatis (Sv3rige) — the sv3rige channel (2011–2025) and the
current Goatis channel (2019–2026). This is the dataset behind
goatis.net, a searchable archive in the style of
aajonus.net.
What makes it more than raw ASR
Every video was processed with speaker identification, not just
transcription. He mostly reacts to other people's videos, so a… See the full description on the dataset page: https://huggingface.co/datasets/exoarbuus/goatis-transcripts.GOAI-MooreSpeechCorpora
Moore Speech Corpora: A Cleaned, Denoised Audio-Text Dataset for Mooré TTS and ASR
The Moore Speech Corpora is a collection of aligned audio and text in Mooré, gathered from publicly available sources. This unified corpus is curated for research and academic purposes in low-resource speech and language processing, especially for text-to-speech (TTS) and automatic speech recognition (ASR) in the Mooré language (ISO 639-3: mos).
Mooré is under-represented in current speech corpora and… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/GOAI-MooreSpeechCorpora.goai-moore-speech-devinettes
Moore Speech Devinettes: A Spoken Riddle Dataset in Mooré
The Moore Speech Devinettes dataset is a spoken collection of traditional Mooré riddles, created for academic and research use in low-resource speech and language technologies.
It is designed to support work in text-to-speech (TTS), automatic speech recognition (ASR), and oral tradition modeling for the Mooré language (ISO 639-3: mos).
[!NOTE]
⚠️ Access is gated. To request access, please read the policy below.
🚩 TLDR: For… See the full description on the dataset page: https://huggingface.co/datasets/goaicorp/goai-moore-speech-devinettes.
