CoolFace
Datasetpublic

gheero-Leyu/leyu-amharic-addis-ababa-dialect

Leyu Amharic - Addis Ababa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-addis-ababa-dialect.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
1likes142downloads
Dataset Card

Leyu Amharic - Addis Ababa Dialect Speech Corpus

Dataset Description

This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and prosodic characteristics unique to the Wello region.

All recordings were collected in real-world environments using mobile devices, introducing natural acoustic variability that improves model generalization. Each audio–text pair underwent manual review to ensure transcript accuracy and audio clarity.

Dataset Summary

Metadata FieldValue
LanguageAmharic (am-ET)
DialectAddis Ababa
Audio Format.wav
Total Hours123.71 Hours
Recording EnvironmentMobile / Crowdsourced (Verified)

Key Statistics

  • —Total Duration: 123.71 Hours
  • —Speaker Count: 86 Speakers (Male: 46, Female: 40)
  • —Data Quality: Mobile-recorded and manually verified for transcript alignment.

Data Collection & Quality Assurance

  • —Recorded by contributors using mobile devices in real-world environments.
  • —Manually reviewed to ensure transcript alignment and audio clarity.

Data Fields

Each sample contains:

  • —text (string): transcript
  • —audio (Audio): waveform/audio file
  • —dialect (string): Addis Ababa
  • —speaker_id (string): anonymized speaker identifier
  • —gender (string): male / female / unknown

iCog Blogs