amharic
Datasets
All datasets matching “amharic”amharic-speech
Dataset.ET Amharic Speech — v0.2.0
51.547 hours · 16,866 clips · 493 speakers · 15,443 distinct prompts
Dataset Summary
Read speech in Amharic, crowdsourced from volunteer contributors in Ethiopia
through a Telegram bot, peer-validated by other contributors, and screened
acoustically before release. Amharic has very little open speech data; this
corpus exists to change that.
Contributors read a displayed prompt aloud, other contributors listen and vote on
whether… See the full description on the dataset page: https://huggingface.co/datasets/snapwre/amharic-speech.amharic-asr-benchmark
Amharic ASR Benchmark
An evaluation of open speech recognition models for Amharic, on a test set with
certain labels and honest statistics.
16 models. 1,548 clips. 4.72 hours. Every hypothesis published.
Published by Dataset.ET.
Read this table first
Round 1 of this benchmark rested on a single clean claim: every model predated
our dataset, so none could have trained on it. That claim no longer holds.
Models trained on snapwre/amharic-speech now exist, and others… See the full description on the dataset page: https://huggingface.co/datasets/snapwre/amharic-asr-benchmark.AmharicCLIP-annotation
AmharicCLIP Annotation Dataset
69,629 images organized by category for Amharic caption annotation.
Structure
images/
animals10/ 18,644 images — 10 animal classes
cat/
dog/
horse/ ...
intel/ 11,998 images — 6 scene classes
forest/
mountain/ ...
fruits360/ 38,987 images — 131 fruit classes
apple/
banana/ ...
Image URL Format… See the full description on the dataset page: https://huggingface.co/datasets/CLIPAMharic/AmharicCLIP-annotation.amharic-tts-benchmark
Amharic TTS Benchmark
Seven text-to-speech systems and the original human recordings, evaluated on 100
Amharic prompts from three open datasets. Run date 2026-08-12.
Published results: addisassistant.com/benchmarks
Reproduce the CER/WER results
python score.py
No arguments. It reads data/judge_rows.jsonl, recomputes every character and
word edit count from the transcripts and writes data/summary.json.
This covers the CER/WER results only. Listening scores… See the full description on the dataset page: https://huggingface.co/datasets/addisai/amharic-tts-benchmark.leyu-amharic-wello-dialect
Leyu Amharic - Wello Dialect Speech Corpus
Dataset Description
This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Wello dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-wello-dialect.waxal-amharic-combined
