datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
waxal-pseudo
WAXAL Pseudo-Labels (3-model agreement)
PRIVATE working artifact for the Google WAXAL ASR Challenge — not for redistribution.
High-confidence pseudo-labels for the WAXAL unlabeled split, produced by a 3-model agreement cascade:
a clip is kept only when the fine-tuned champion (w2v-BERT-2.0 CTC) and XLS-R-300m agree
(CER ≤ 0.12), and the fine-tuned Omnilingual-ASR-300M independently confirms the champion transcript
(CER ≤ 0.22). omni is architecturally diverse (different… See the full description on the dataset page: https://huggingface.co/datasets/anyantudre/waxal-pseudo.waxal-linsna
WAXAL Phase-2 Lingala / Shona training corpus (derived)
This corpus was built for the Google WAXAL ASR Challenge on Zindi.
This repository is the exact training corpus behind our Google WAXAL ASR Challenge (Phase 2)
submission: the TSV manifests plus the derived 16 kHz mono FLAC audio that our training configs read.
It exists so that the whole recipe can be rebuilt and audited from one place. The manifests below,
together with the google/WaxalNLP train and validation splits read… See the full description on the dataset page: https://huggingface.co/datasets/anyantudre/waxal-linsna.moore-speech-devinettes
Moore Speech Devinettes: A Spoken Riddle Dataset in Mooré
The Moore Speech Devinettes dataset is a spoken collection of traditional Mooré riddles, created for academic and research use in low-resource speech and language technologies.
It is designed to support work in text-to-speech (TTS), automatic speech recognition (ASR), and oral tradition modeling for the Mooré language (ISO 639-3: mos).
[!NOTE]
⚠️ Access is gated. To request access, please read the policy below.
🚩 TLDR: For… See the full description on the dataset page: https://huggingface.co/datasets/anyantudre/moore-speech-devinettes.MooreSpeechCorpora
Moore Speech Corpora: A Cleaned, Denoised Audio-Text Dataset for Mooré TTS and ASR
The Moore Speech Corpora is a collection of aligned audio and text in Mooré, gathered from publicly available sources. This unified corpus is curated for research and academic purposes in low-resource speech and language processing, especially for text-to-speech (TTS) and automatic speech recognition (ASR) in the Mooré language (ISO 639-3: mos).
Mooré is under-represented in current speech corpora and… See the full description on the dataset page: https://huggingface.co/datasets/anyantudre/MooreSpeechCorpora.moore-speech-contes
Moore Speech Contes: A Spoken Corpus of Traditional Mooré Stories
The Moore Speech Contes dataset is a collection of spoken folk stories (contes) in Mooré, designed for research and academic purposes in low-resource speech and language processing.
It is intended primarily for academic or research purposes in text-to-speech (TTS) and automatic speech recognition (ASR) for Mooré language (ISO 639-3: mos).
[!NOTE]
⚠️ Access is gated. To request access, please read the policy below.🚩… See the full description on the dataset page: https://huggingface.co/datasets/anyantudre/moore-speech-contes.moore-speech-proverbes
Moore Speech Proverbs: A Parallel Audio-Text Corpus for Mooré and French
The Moore Speech Proverbs dataset is a bilingual audio-text corpus of traditional proverbs in Mooré and French, designed for research and academic purposes in low-resource speech and language processing.
It is intended primarily for academic or research purposes in text-to-speech (TTS) and automatic speech recognition (ASR) for Mooré language.
[!NOTE]
⚠️ Access is gated. To request access, please read the… See the full description on the dataset page: https://huggingface.co/datasets/anyantudre/moore-speech-proverbes.
