datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
afaan-oromoo-speech
Dataset.ET Afaan Oromoo Speech — v0.1.0
9.843 hours · 3,594 clips · 74 speakers · 3,283 distinct prompts
Dataset Summary
Read speech in Afaan Oromoo, crowdsourced from volunteer contributors in Ethiopia
through a Telegram bot, peer-validated by other contributors, and screened
acoustically before release. Afaan Oromoo has very little open speech data; this
corpus exists to change that.
Contributors read a displayed prompt aloud, other contributors listen and vote… See the full description on the dataset page: https://huggingface.co/datasets/snapwre/afaan-oromoo-speech.leyu-oromo-speech-corpus-2026
Leyu Afaan Oromo Speech Corpus 2026
Official speech dataset submission for the Leyu Data Collection Competition 2026.
Organization & Team
Hugging Face Org: SoundWaveET
Dataset Repo: SoundWaveET/leyu-oromo-speech-corpus-2026
