datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leyu-amharic-gonder-dialect
Leyu Amharic - Gonder Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gonder dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gonder-dialect.leyu-amharic-gojjam-dialect
Leyu Amharic - Gojjam Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gojjam dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gojjam-dialect.leyu-amharic-wello-dialect
Leyu Amharic - Wello Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Wello dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-wello-dialect.leyu-amharic-addis-ababa-dialect
Leyu Amharic - Addis Ababa Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Addis Ababa dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-addis-ababa-dialect.leyu-amharic-shewa-dialect
Leyu Amharic - Shewa Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Shewa dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-shewa-dialect.fleurs-ethiopian-v2
FLEURS — Ethiopian Languages
This dataset is a restructured v2 conversion of the Google FLEURS dataset for two Ethiopian languages: Amharic (am_et) and Oromo (om_et).
Subsets
Subset
Language
ISO 639-2
Train
Dev
Test
amh
Amharic
amh
3,163
223
516
orm
Oromo
orm
1,701
19
41
Splits
Split
Description
train
Training split
dev
Development split (renamed from validation in original FLEURS)
test
Test split
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/fleurs-ethiopian-v2.alffa-amharic
ALFFA Amharic Speech Corpus
Read speech corpus for Amharic (አማርኛ) automatic speech recognition, converted to HuggingFace Datasets format from the original ALFFA project.
Dataset Structure
{
'audio': Audio(sampling_rate=16000),
'utterance_id': 'tr_10000_tr097082',
'transcript': 'ይህ አማርኛ ጽሑፍ ነው',
'speaker_id': '097',
'split': 'train'
}
Usage
from datasets import load_dataset
dataset = load_dataset("hadamard-2/alffa-amharic")
# Access… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/alffa-amharic.sagalee
Sagalee — Oromo ASR Dataset
This is a verbatim upload of the Sagalee dataset, an open-source automatic speech recognition dataset for the Oromo language, originally released by Turi Abu et al. and accepted at ICASSP 2025.
License notice: This dataset is released under CC BY-NC 4.0 — use for commercial purposes is not permitted.
Subset
Single subset orm (Oromo, ISO 639-2). No subset needed for single-language datasets.
Splits
Split
Speakers… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/sagalee.alffa-amharic-v2
ALFFA Amharic Speech Corpus (v2)
Read speech corpus for Amharic (አማርኛ) automatic speech recognition. Converted from the original ALFFA project and restructured to match the google/waxalnlp schema for interoperability.
This is a restructured version of hadamard-2/alffa-amharic.
Changes from v1
utterance_id renamed to id
transcript renamed to transcription
speaker_id set to "unknown" — the original ALFFA Kaldi files shipped with utt2spk mapping each utterance to itself… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/alffa-amharic-v2.common-voice-24-ethiopian-v2
Common Voice 24.0 — Ethiopian Languages (v2)
This is a cleaned and restructured version of hadamard-2/common-voice-24-ethiopian, which is a verbatim archival upload of the Mozilla Common Voice 24.0 scripted speech data for Amharic and Tigrinya. This version conforms to the Waxal ASR schema.
Subsets
Subset
Language
ISO 639-2
Clips
Validated Hours
Speakers
amh
Amharic
amh
1,045
1.82h
46
tir
Tigrinya
tir
69
0.10h
16
Splits
Split… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/common-voice-24-ethiopian-v2.sagalee-v2
Sagalee — Oromo ASR Dataset (v2)
This is a restructured v2 version of hadamard-2/sagalee, which is a verbatim upload of the Sagalee dataset — an open-source ASR dataset for the Oromo language, originally released by Turi Abu et al. and accepted at ICASSP 2025.
License notice: This dataset is released under CC BY-NC 4.0 — use for commercial purposes is not permitted.
Subset
Single subset orm (Oromo, ISO 639-2). No subset needed for single-language datasets.… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/sagalee-v2.common-voice-24-ethiopian
Common Voice 24.0 — Ethiopian Languages
This dataset is a verbatim archival upload of the Mozilla Common Voice 24.0 scripted speech data for two Ethiopian languages: Amharic (am) and Tigrinya (ti), sourced from the Mozilla Data Collective.
Subsets
Subset
Language
Code
Clips
Total Hours
Validated Hours
Speakers
amharic
Amharic
am
1,632
2.85h
1.82h
46
tigrinya
Tigrinya
ti
451
0.65h
0.10h
16
Splits
Each subset contains the following splits:… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/common-voice-24-ethiopian.hadamard-2_algorithmic_improvementhadamard-1_finetunehadamard-3_novel_breakthrough
