amh
Datasets
All datasets matching “amh”amharic-speech
Dataset.ET Amharic Speech — v0.2.0
51.547 hours · 16,866 clips · 493 speakers · 15,443 distinct prompts
Dataset Summary
Read speech in Amharic, crowdsourced from volunteer contributors in Ethiopia
through a Telegram bot, peer-validated by other contributors, and screened
acoustically before release. Amharic has very little open speech data; this
corpus exists to change that.
Contributors read a displayed prompt aloud, other contributors listen and vote on
whether… See the full description on the dataset page: https://huggingface.co/datasets/snapwre/amharic-speech.wxl_amhAmharicCLIP-annotation
AmharicCLIP Annotation Dataset
69,629 images organized by category for Amharic caption annotation.
Structure
images/
animals10/ 18,644 images — 10 animal classes
cat/
dog/
horse/ ...
intel/ 11,998 images — 6 scene classes
forest/
mountain/ ...
fruits360/ 38,987 images — 131 fruit classes
apple/
banana/ ...
Image URL Format… See the full description on the dataset page: https://huggingface.co/datasets/CLIPAMharic/AmharicCLIP-annotation.amharic-tts-benchmark
Amharic TTS Benchmark
Seven text-to-speech systems and the original human recordings, evaluated on 100
Amharic prompts from three open datasets. Run date 2026-08-12.
Published results: addisassistant.com/benchmarks
Reproduce the CER/WER results
python score.py
No arguments. It reads data/judge_rows.jsonl, recomputes every character and
word edit count from the transcripts and writes data/summary.json.
This covers the CER/WER results only. Listening scores… See the full description on the dataset page: https://huggingface.co/datasets/addisai/amharic-tts-benchmark.amharic-asr-benchmark
Amharic ASR Benchmark
An evaluation of open speech recognition models for Amharic, on a test set with
certain labels and honest statistics.
16 models. 1,548 clips. 4.72 hours. Every hypothesis published.
Published by Dataset.ET.
Read this table first
Round 1 of this benchmark rested on a single clean claim: every model predated
our dataset, so none could have trained on it. That claim no longer holds.
Models trained on snapwre/amharic-speech now exist, and others… See the full description on the dataset page: https://huggingface.co/datasets/snapwre/amharic-asr-benchmark.waxal-amharic-combined
