datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
amharic-speech
Dataset.ET Amharic Speech — v0.2.0
51.547 hours · 16,866 clips · 493 speakers · 15,443 distinct prompts
Dataset Summary
Read speech in Amharic, crowdsourced from volunteer contributors in Ethiopia
through a Telegram bot, peer-validated by other contributors, and screened
acoustically before release. Amharic has very little open speech data; this
corpus exists to change that.
Contributors read a displayed prompt aloud, other contributors listen and vote on
whether… See the full description on the dataset page: https://huggingface.co/datasets/snapwre/amharic-speech.amharic-tts-benchmark
Amharic TTS Benchmark
Seven text-to-speech systems and the original human recordings, evaluated on 100
Amharic prompts from three open datasets. Run date 2026-08-12.
Published results: addisassistant.com/benchmarks
Reproduce the CER/WER results
python score.py
No arguments. It reads data/judge_rows.jsonl, recomputes every character and
word edit count from the transcripts and writes data/summary.json.
This covers the CER/WER results only. Listening scores… See the full description on the dataset page: https://huggingface.co/datasets/addisai/amharic-tts-benchmark.waxal-amharic-combinedleyu-amharic-wello-dialect
Leyu Amharic - Wello Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Wello dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-wello-dialect.leyu-amharic-addis-ababa-dialect
Leyu Amharic - Addis Ababa Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-addis-ababa-dialect.amharic-speech-datasetamharic-pretraining-corpusAmharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining tasks. It consists of diverse text sources, including news articles, books, social media posts, government documents, and web content, all written in Amharic.
You can load the dataset as follows
from datasets import load_dataset
ds = load_dataset("yordanoswuletaw/amharic-pretraining-corpus")
amharic-speech-expandedamharic-audio
🎵 Amharic Bible Audio Dataset
📋 Dataset Description
This dataset contains 59K audio chunks derived from Amharic Bible readings, split into 5-second segments for optimal training of speech models.
Audio Specifications
Format: WAV (16-bit PCM)
Sample Rate: 24 kHz
Duration: 5 seconds per chunk
Total Hours: ~82.5 hours
🚀 Usage
Load with HuggingFace Datasets
from datasets import load_dataset
# Load the dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/NaolBM/amharic-audio.leyu-amharic-shewa-dialect
Leyu Amharic - Shewa Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Shewa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-shewa-dialect.english-amharic_sentence-pairs_mt560
English-Amharic Parallel Dataset
This dataset contains parallel sentences in English and Amharic (Ethiopia).
Dataset Information
Language Pair: English ↔ Amharic
Language Code: amh
Country: Ethiopia
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this dataset, please cite the… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-amharic_sentence-pairs_mt560.leyu-amharic-gonder-dialect
Leyu Amharic - Gonder Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gonder dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gonder-dialect.amharic-combined-corpusleyu-amharic-gojjam-dialect
Leyu Amharic - Gojjam Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gojjam dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gojjam-dialect.amharichatespeechranlp
Introduction
The Amharic Hate Speech data is collected using the Twitter API spanning from October 1, 2020 - November 30, 2022, considering the socio-political dynamics of Ethiopia in Twitter space. We used WebAnno tool for data annotation; each tweet is annotated by two native speakers and curated by one more experienced adjudicator to determine the gold labels. A total of 15.1k tweets consisting of three class labels namely: Hate, Offensive and Normal are presented. Read our… See the full description on the dataset page: https://huggingface.co/datasets/uhhlt/amharichatespeechranlp.leyu-amharic-wello-dialect
Leyu Amharic - Wello Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Wello dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-wello-dialect.Kuzi-Amharic-Uncensored-Datasetleyu-amharic-gonder-dialect
Leyu Amharic - Gonder Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gonnder dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-gonder-dialect.leyu-amharic-gojjam-dialect
Leyu Amharic - Gojjam Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-gojjam-dialect.leyu-amharic-shewa-dialect
Leyu Amharic - Shewa Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Shewa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-shewa-dialect.leyu-amharic-addis-ababa-dialect
Leyu Amharic - Addis Ababa Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Addis Ababa dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-addis-ababa-dialect.amharic-sentences-corpus
Amharic Sentences Corpus V1.0
Source: Telegram
This 1.6 million Amharic sentences corpus reflects current Amharic usage as of December 20, 2025, and is designed for anyone interested in:
Training Amharic-based LLMs
Fine-tuning NLP models
Building search, summarization, or generative systems in Amharic
The dataset is heavily cleaned and normalized, but like any serious LLM dataset, it still needs proper tokenization for pre-training.
I recommend using an… See the full description on the dataset page: https://huggingface.co/datasets/a3xrfgb/amharic-sentences-corpus.leyu-amharic-shewa-dialect
Leyu Amharic - Shewa Dialect Speech Corpus
Dataset Description
A parallel speech corpus of audio recordings paired with their transcripts, focused on the Shewa dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity.
This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-shewa-dialect.amharic-cpt-corpus-v16-balanced
📊 Dataset Token Distribution & Statistics
This dataset has been cleaned and tokenized using the abdukuzi45/qwen3.5-4b-amharic-v4 tokenizer.
Category (Source)
Token Count
Percentage
Row Count
🇪🇹 Amharic
1,159,395,439
37.82%
1,244,733
💻 Code
724,855,298
23.65%
764,410
📐 Math
680,130,234
22.19%
728,000
🇬🇧 English
500,648,252
16.34%
559,054
Total
3,065,029,223
100.0%
3,296,197
Key Highlights
Total Tokens: ~3.065 Billion Tokens
Primary… See the full description on the dataset page: https://huggingface.co/datasets/abdukuzi45/amharic-cpt-corpus-v16-balanced.leyu-amharic-gojjam-dialect
Leyu Amharic - Gojjam Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-gojjam-dialect.Amharic_Instruction_dataset
SFT-Data for Walia-LLM: Enhancing Amharic-LLaMA by Integrating Task-Specific and Generative Datasets
Dataset Summary
The Walia dataset is designed to enhance large language models for the Amharic language by:
Converting existing task-specific datasets (e.g., sentiment analysis, QA, NER) into instruction format.
Creating new generative datasets (e.g., poem generation, religious lyrics, story generation).
Translating English instruction datasets (e.g., Alpaca, Dolly) into… See the full description on the dataset page: https://huggingface.co/datasets/EthioNLP/Amharic_Instruction_dataset.amharic-sentences-corpus
Amharic Sentences Corpus
This dataset is compiled by Yimam et al. (2021) at LT Group, University of Hamburg, Germany. It comprises a collection of 6.4 million Amharic sentences intended for use in language model pretraining.
Source
GitHub https://github.com/uhh-lt/ethiopicmodels
Dataset: https://data.mendeley.com/datasets/dtywyf3sth/1
Paper: https://www.mdpi.com/1999-5903/13/11/275
For citing this dataset, please use the following:
@Article{fi13110275,
AUTHOR = {Yimam… See the full description on the dataset page: https://huggingface.co/datasets/rasyosef/amharic-sentences-corpus.leyu-amharic-addis-ababa-dialect
Leyu Amharic - Addis Ababa Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-addis-ababa-dialect.leyu-amharic-wello-dialect
Leyu Amharic - Wello Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Wello dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-wello-dialect.amharic-speech-dataset-110HRS-V21
