gheero-Leyu/leyu-amharic-gojjam-dialect
Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-gojjam-dialect.
Leyu Amharic - Gojjam Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and prosodic characteristics that are essential for building dialect-aware and robust speech systems.
All recordings were collected in real-world environments using mobile devices, introducing natural acoustic variability that improves model generalization. Each audio–text pair underwent manual review to ensure transcript accuracy and audio clarity. By emphasizing demographic diversity and regional authenticity, the dataset contributes to the advancement of inclusive and representative speech technologies for low-resource African languages.
Dataset Summary
Key Statistics
- Total Duration: 95.77 Hours
- Speaker Count: 69 Speakers (Male: 41, Female: 28)
- Data Quality: Mobile-recorded and manually verified for transcript alignment.
Data Collection & Quality Assurance
- Recorded by contributors using mobile devices in real-world environments.
- Manually reviewed to ensure transcript alignment and audio clarity.
Data Fields
Each sample contains:
text(string): transcriptaudio(Audio): waveform/audio filedialect(string):gojjamspeaker_id(string): anonymized speaker identifiergender(string):male/female/unknown
gheero Blogs
- Leyu: Crowdsourcing Datasets for Ethiopian Languages
- Dialects & Socioeconomics: Shaping Inclusive Language Models
- Progress of Natural Language Processing (NLP) for Ethiopian Languages – Part One
- Progress of Natural Language Processing (NLP) for Ethiopian Languages – Part Two
- Data Collection with Purpose: The Leyu Approach
