leyu-amharic/leyu-amharic-wello-dialect
Leyu Amharic - Wello Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Wello dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-wello-dialect.
Leyu Amharic - Wello Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Wello dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and prosodic characteristics unique to the Wello region.
All recordings were collected in real-world environments using mobile devices, introducing natural acoustic variability that improves model generalization. Each audio–text pair underwent manual review to ensure transcript accuracy and audio clarity.
Dataset Summary
Key Statistics
- Total Duration: 114.7 Hours
- Speaker Count: 68 Speakers (Male:27, Female: 41)
- Data Quality: Mobile-recorded and manually verified for transcript alignment.
Data Collection & Quality Assurance
- Recorded by contributors using mobile devices in real-world environments.
- Manually reviewed to ensure transcript alignment and audio clarity.
Data Fields
Each sample contains:
text(string): transcriptaudio(Audio): waveform/audio filedialect(string):wellospeaker_id(string): anonymized speaker identifiergender(string):male/female/unknown
iCog Blogs
- Leyu: Crowdsourcing Datasets for Ethiopian Languages
- Dialects & Socioeconomics: Shaping Inclusive Language Models
- Progress of Natural Language Processing (NLP) for Ethiopian Languages – Part One
- Progress of Natural Language Processing (NLP) for Ethiopian Languages – Part Two
- Data Collection with Purpose: The Leyu Approach
