CoolFace
Datasetpublic

gheero-Leyu/leyu-amharic-gojjam-dialect

Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-gojjam-dialect.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes193downloads
Dataset Card

Leyu Amharic - Gojjam Dialect Speech Corpus

Dataset Description

This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and prosodic characteristics that are essential for building dialect-aware and robust speech systems.

All recordings were collected in real-world environments using mobile devices, introducing natural acoustic variability that improves model generalization. Each audio–text pair underwent manual review to ensure transcript accuracy and audio clarity. By emphasizing demographic diversity and regional authenticity, the dataset contributes to the advancement of inclusive and representative speech technologies for low-resource African languages.

Dataset Summary

Metadata FieldValue
LanguageAmharic (am-ET)
DialectGojjam
Audio Format.wav
Total Hours95.77 Hours
Recording EnvironmentMobile / Crowdsourced (Verified)

Key Statistics

  • —Total Duration: 95.77 Hours
  • —Speaker Count: 69 Speakers (Male: 41, Female: 28)
  • —Data Quality: Mobile-recorded and manually verified for transcript alignment.

Data Collection & Quality Assurance

  • —Recorded by contributors using mobile devices in real-world environments.
  • —Manually reviewed to ensure transcript alignment and audio clarity.

Data Fields

Each sample contains:

  • —text (string): transcript
  • —audio (Audio): waveform/audio file
  • —dialect (string): gojjam
  • —speaker_id (string): anonymized speaker identifier
  • —gender (string): male / female / unknown

gheero Blogs