CoolFace
Datasetpublic

gheero-Leyu/leyu-amharic-gonder-dialect

Leyu Amharic - Gonder Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gonnder dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-gonder-dialect.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
1likes43downloads
Dataset Card

Leyu Amharic - Gonder Dialect Speech Corpus

Dataset Description

This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gonnder dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and prosodic characteristics unique to the Wello region.

All recordings were collected in real-world environments using mobile devices, introducing natural acoustic variability that improves model generalization. Each audio–text pair underwent manual review to ensure transcript accuracy and audio clarity.

Dataset Summary

Metadata FieldValue
LanguageAmharic (am-ET)
DialectGonder
Audio Format.wav
Total Hours112.83 Hours
Recording EnvironmentMobile / Crowdsourced (Verified)

Key Statistics

  • —Total Duration: 112.83 Hours
  • —Speaker Count: 67 Speakers (Male: 23, Female: 44)
  • —Data Quality: Mobile-recorded and manually verified for transcript alignment.

Data Collection & Quality Assurance

  • —Recorded by contributors using mobile devices in real-world environments.
  • —Manually reviewed to ensure transcript alignment and audio clarity.

Data Fields

Each sample contains:

  • —text (string): transcript
  • —audio (Audio): waveform/audio file
  • —dialect (string): gonder
  • —speaker_id (string): anonymized speaker identifier
  • —gender (string): male / female / unknown

iCog Blogs