datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
amharic-speech
Dataset.ET Amharic Speech — v0.2.0
51.547 hours · 16,866 clips · 493 speakers · 15,443 distinct prompts
Dataset Summary
Read speech in Amharic, crowdsourced from volunteer contributors in Ethiopia
through a Telegram bot, peer-validated by other contributors, and screened
acoustically before release. Amharic has very little open speech data; this
corpus exists to change that.
Contributors read a displayed prompt aloud, other contributors listen and vote on
whether… See the full description on the dataset page: https://huggingface.co/datasets/snapwre/amharic-speech.leyu-amharic-wello-dialect
Leyu Amharic - Wello Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Wello dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-wello-dialect.leyu-amharic-addis-ababa-dialect
Leyu Amharic - Addis Ababa Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-addis-ababa-dialect.leyu-amharic-shewa-dialect
Leyu Amharic - Shewa Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Shewa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-shewa-dialect.leyu-amharic-gonder-dialect
Leyu Amharic - Gonder Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gonder dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gonder-dialect.leyu-amharic-gojjam-dialect
Leyu Amharic - Gojjam Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gojjam dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gojjam-dialect.leyu-amharic-wello-dialect
Leyu Amharic - Wello Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Wello dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-wello-dialect.leyu-amharic-gonder-dialect
Leyu Amharic - Gonder Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gonnder dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-gonder-dialect.leyu-amharic-gojjam-dialect
Leyu Amharic - Gojjam Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-gojjam-dialect.leyu-amharic-shewa-dialect
Leyu Amharic - Shewa Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Shewa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-shewa-dialect.leyu-amharic-addis-ababa-dialect
Leyu Amharic - Addis Ababa Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Addis Ababa dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-addis-ababa-dialect.leyu-amharic-shewa-dialect
Leyu Amharic - Shewa Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Shewa dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-shewa-dialect.leyu-amharic-gojjam-dialect
Leyu Amharic - Gojjam Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-gojjam-dialect.leyu-amharic-addis-ababa-dialect
Leyu Amharic - Addis Ababa Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-addis-ababa-dialect.amharic-bdu-asr
Amharic ASR Dataset
Dataset Summary
This dataset is a processed version of the original BDU-speech dataset by Yohannes A. Ejigu.
It contains paired Amharic speech audio and transcriptions, structured for use in automatic speech recognition (ASR) research and model training.
Audio files are decoded as mono. Sampling rates may vary across files.
Dataset Info
DatasetDict({
train: Dataset({
features: ['audio', 'sentence'],
num_rows: 32901… See the full description on the dataset page: https://huggingface.co/datasets/chappM/amharic-bdu-asr.leyu-amharic-wello-dialect
Leyu Amharic - Wello Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Wello dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-wello-dialect.alffa-amharic
ALFFA Amharic Speech Corpus
Read speech corpus for Amharic (አማርኛ) automatic speech recognition, converted to HuggingFace Datasets format from the original ALFFA project.
Dataset Structure
{
'audio': Audio(sampling_rate=16000),
'utterance_id': 'tr_10000_tr097082',
'transcript': 'ይህ አማርኛ ጽሑፍ ነው',
'speaker_id': '097',
'split': 'train'
}
Usage
from datasets import load_dataset
dataset = load_dataset("hadamard-2/alffa-amharic")
# Access… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/alffa-amharic.leyu-amharic-gonder-dialect
Leyu Amharic - Gonder Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gonnder dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-gonder-dialect.amharic-speech-transcribed
Amharic Spontaneous Speech, Transcribed — Silencio
Spontaneous Amharic with human-validated transcription in Ge'ez script and word-level alignment. 45 clips from 23 distinct native speakers in Ethiopia, most of them from Addis Ababa, with a Gojjam group. Transcripts are written in Fidel (the Ethiopic script), not romanised.
Hours
0.48
Clips
45
Speakers
23
Countries
1
Speaker origin regions
3
Native Amharic speakers (declared)
23 of 23 (45 clips)
Audio
48… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/amharic-speech-transcribed.amharic-asr-merged
amharic-asr-merged
A merged Amharic speech-recognition dataset, combining and deduplicating:
badrex/amharic-speech
chappM/amharic-bdu-asr
beimnet777/amharic-asr
snapwre/amharic-speech
Processing
Standardized to audio (16kHz mono) and text columns, with a source column tracking origin
Unicode NFC-normalized transcripts, empty transcripts dropped
Exact-duplicate transcripts removed
Re-split into train (90%) / validation (5%) / test (5%), ignoring original source… See the full description on the dataset page: https://huggingface.co/datasets/Harbidel/amharic-asr-merged.alffa-amharic-v2
ALFFA Amharic Speech Corpus (v2)
Read speech corpus for Amharic (አማርኛ) automatic speech recognition. Converted from the original ALFFA project and restructured to match the google/waxalnlp schema for interoperability.
This is a restructured version of hadamard-2/alffa-amharic.
Changes from v1
utterance_id renamed to id
transcript renamed to transcription
speaker_id set to "unknown" — the original ALFFA Kaldi files shipped with utt2spk mapping each utterance to itself… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/alffa-amharic-v2.Lugha
Lugha – Hausa Speech Dataset
Crowd-sourced Hausa voice recordings collected via the Lugha mobile app.
Structure
audio/sample_XXXX.m4a – raw audio (m4a)
metadata.jsonl – one JSON object per recording
Fields
field
description
audio
relative path to the audio file
text
prompt sentence read by the speaker
language
spoken language
state
Nigerian state of the speaker
lga
Local Government Area
accent
self-reported accent
age_range… See the full description on the dataset page: https://huggingface.co/datasets/Amhaztech/Lugha.leyu-amharic-shewa-dialect
Leyu Amharic - Shewa Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Shewa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/birhanu23/leyu-amharic-shewa-dialect.
