CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01snapwre /amharic-speech Dataset.ET Amharic Speech — v0.2.0 51.547 hours · 16,866 clips · 493 speakers · 15,443 distinct prompts Dataset Summary Read speech in Amharic, crowdsourced from volunteer contributors in Ethiopia through a Telegram bot, peer-validated by other contributors, and screened acoustically before release. Amharic has very little open speech data; this corpus exists to change that. Contributors read a displayed prompt aloud, other contributors listen and vote on whether… See the full description on the dataset page: https://huggingface.co/datasets/snapwre/amharic-speech.audioautomatic-speech-recognition10K<n<100K25 likes1.2k downloads23d agoHugging Face02addisai /amharic-tts-benchmark Amharic TTS Benchmark Seven text-to-speech systems and the original human recordings, evaluated on 100 Amharic prompts from three open datasets. Run date 2026-08-12. Published results: addisassistant.com/benchmarks Reproduce the CER/WER results python score.py No arguments. It reads data/judge_rows.jsonl, recomputes every character and word edit count from the transcripts and writes data/summary.json. This covers the CER/WER results only. Listening scores… See the full description on the dataset page: https://huggingface.co/datasets/addisai/amharic-tts-benchmark.audiotext-to-speechn<1K0 likes639 downloads1mo agoHugging Face03SaarAI /waxal-amharic-combinedaudio100K<n<1M0 likes496 downloads1mo agoHugging Face04leyu-amharic /leyu-amharic-wello-dialect Leyu Amharic - Wello Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Wello dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-wello-dialect.audioautomatic-speech-recognition10K<n<100K1 likes435 downloads2mo agoHugging Face05leyu-amharic /leyu-amharic-addis-ababa-dialect Leyu Amharic - Addis Ababa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-addis-ababa-dialect.audioautomatic-speech-recognition1K<n<10K1 likes374 downloads2mo agoHugging Face06shunyalabs /amharic-speech-datasetaudio10K<n<100K2 likes341 downloads1y agoHugging Face07yordanoswuletaw /amharic-pretraining-corpusAmharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining tasks. It consists of diverse text sources, including news articles, books, social media posts, government documents, and web content, all written in Amharic. You can load the dataset as follows from datasets import load_dataset ds = load_dataset("yordanoswuletaw/amharic-pretraining-corpus") texttext-generation100M<n<1B4 likes296 downloads2y agoHugging Face08israel /amharic-speech-expandedaudio10K<n<100K0 likes292 downloads3mo agoHugging Face09NaolBM /amharic-audio 🎵 Amharic Bible Audio Dataset 📋 Dataset Description This dataset contains 59K audio chunks derived from Amharic Bible readings, split into 5-second segments for optimal training of speech models. Audio Specifications Format: WAV (16-bit PCM) Sample Rate: 24 kHz Duration: 5 seconds per chunk Total Hours: ~82.5 hours 🚀 Usage Load with HuggingFace Datasets from datasets import load_dataset # Load the dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/NaolBM/amharic-audio.audio10K<n<100K0 likes275 downloads7mo agoHugging Face10leyu-amharic /leyu-amharic-shewa-dialect Leyu Amharic - Shewa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Shewa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-shewa-dialect.audioautomatic-speech-recognition1K<n<10K1 likes252 downloads2mo agoHugging Face11michsethowusu /english-amharic_sentence-pairs_mt560 English-Amharic Parallel Dataset This dataset contains parallel sentences in English and Amharic (Ethiopia). Dataset Information Language Pair: English ↔ Amharic Language Code: amh Country: Ethiopia Original Source: OPUS MT560 Dataset Dataset Structure The dataset contains parallel sentences that can be used for: Machine translation training Cross-lingual NLP tasks Language model fine-tuning Citation If you use this dataset, please cite the… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-amharic_sentence-pairs_mt560.text100K<n<1M2 likes232 downloads1y agoHugging Face12hadamard-2 /leyu-amharic-gonder-dialect Leyu Amharic - Gonder Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gonder dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gonder-dialect.audioautomatic-speech-recognition10K<n<100K0 likes221 downloads10d agoHugging Face13b1n1yam /amharic-combined-corpustext10M<n<100M0 likes216 downloads10mo agoHugging Face14hadamard-2 /leyu-amharic-gojjam-dialect Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Gojjam dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-gojjam-dialect.audioautomatic-speech-recognition10K<n<100K0 likes216 downloads10d agoHugging Face15uhhlt /amharichatespeechranlp Introduction The Amharic Hate Speech data is collected using the Twitter API spanning from October 1, 2020 - November 30, 2022, considering the socio-political dynamics of Ethiopia in Twitter space. We used WebAnno tool for data annotation; each tweet is annotated by two native speakers and curated by one more experienced adjudicator to determine the gold labels. A total of 15.1k tweets consisting of three class labels namely: Hate, Offensive and Normal are presented. Read our… See the full description on the dataset page: https://huggingface.co/datasets/uhhlt/amharichatespeechranlp.texttext-classification10K<n<100K1 likes211 downloads2y agoHugging Face16hadamard-2 /leyu-amharic-wello-dialect Leyu Amharic - Wello Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Wello dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-wello-dialect.audioautomatic-speech-recognition10K<n<100K0 likes205 downloads10d agoHugging Face17abdukuzi45 /Kuzi-Amharic-Uncensored-Datasettext100K<n<1M1 likes201 downloads10d agoHugging Face18leyu-amharic /leyu-amharic-gonder-dialect Leyu Amharic - Gonder Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gonnder dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-gonder-dialect.audioautomatic-speech-recognition10K<n<100K1 likes190 downloads2mo agoHugging Face19gheero-Leyu /leyu-amharic-gojjam-dialect Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-gojjam-dialect.audioautomatic-speech-recognition10K<n<100K0 likes188 downloads2mo agoHugging Face20gheero-Leyu /leyu-amharic-shewa-dialect Leyu Amharic - Shewa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Shewa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-shewa-dialect.audioautomatic-speech-recognition1K<n<10K0 likes184 downloads2mo agoHugging Face21hadamard-2 /leyu-amharic-addis-ababa-dialect Leyu Amharic - Addis Ababa Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Addis Ababa dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-addis-ababa-dialect.audioautomatic-speech-recognition1K<n<10K0 likes178 downloads10d agoHugging Face22a3xrfgb /amharic-sentences-corpus Amharic Sentences Corpus V1.0 Source: Telegram This 1.6 million Amharic sentences corpus reflects current Amharic usage as of December 20, 2025, and is designed for anyone interested in: Training Amharic-based LLMs Fine-tuning NLP models Building search, summarization, or generative systems in Amharic The dataset is heavily cleaned and normalized, but like any serious LLM dataset, it still needs proper tokenization for pre-training. I recommend using an… See the full description on the dataset page: https://huggingface.co/datasets/a3xrfgb/amharic-sentences-corpus.text1M<n<10M1 likes165 downloads7mo agoHugging Face23hadamard-2 /leyu-amharic-shewa-dialect Leyu Amharic - Shewa Dialect Speech Corpus Dataset Description A parallel speech corpus of audio recordings paired with their transcripts, focused on the Shewa dialect of Amharic, for ASR and TTS research. Leyu reports that recordings were collected from contributors on mobile devices in real-world environments, and that each audio–text pair was manually reviewed for transcript accuracy and audio clarity. This repository is a copy of… See the full description on the dataset page: https://huggingface.co/datasets/hadamard-2/leyu-amharic-shewa-dialect.audioautomatic-speech-recognition1K<n<10K0 likes150 downloads10d agoHugging Face24abdukuzi45 /amharic-cpt-corpus-v16-balanced 📊 Dataset Token Distribution & Statistics This dataset has been cleaned and tokenized using the abdukuzi45/qwen3.5-4b-amharic-v4 tokenizer. Category (Source) Token Count Percentage Row Count 🇪🇹 Amharic 1,159,395,439 37.82% 1,244,733 💻 Code 724,855,298 23.65% 764,410 📐 Math 680,130,234 22.19% 728,000 🇬🇧 English 500,648,252 16.34% 559,054 Total 3,065,029,223 100.0% 3,296,197 Key Highlights Total Tokens: ~3.065 Billion Tokens Primary… See the full description on the dataset page: https://huggingface.co/datasets/abdukuzi45/amharic-cpt-corpus-v16-balanced.text1M<n<10M0 likes150 downloads4d agoHugging Face25leyu-amharic /leyu-amharic-gojjam-dialect Leyu Amharic - Gojjam Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Gojjam dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/leyu-amharic/leyu-amharic-gojjam-dialect.audioautomatic-speech-recognition10K<n<100K1 likes149 downloads2mo agoHugging Face26EthioNLP /Amharic_Instruction_dataset SFT-Data for Walia-LLM: Enhancing Amharic-LLaMA by Integrating Task-Specific and Generative Datasets Dataset Summary The Walia dataset is designed to enhance large language models for the Amharic language by: Converting existing task-specific datasets (e.g., sentiment analysis, QA, NER) into instruction format. Creating new generative datasets (e.g., poem generation, religious lyrics, story generation). Translating English instruction datasets (e.g., Alpaca, Dolly) into… See the full description on the dataset page: https://huggingface.co/datasets/EthioNLP/Amharic_Instruction_dataset.text100K<n<1M5 likes134 downloads1y agoHugging Face27rasyosef /amharic-sentences-corpus Amharic Sentences Corpus This dataset is compiled by Yimam et al. (2021) at LT Group, University of Hamburg, Germany. It comprises a collection of 6.4 million Amharic sentences intended for use in language model pretraining. Source GitHub https://github.com/uhh-lt/ethiopicmodels Dataset: https://data.mendeley.com/datasets/dtywyf3sth/1 Paper: https://www.mdpi.com/1999-5903/13/11/275 For citing this dataset, please use the following: @Article{fi13110275, AUTHOR = {Yimam… See the full description on the dataset page: https://huggingface.co/datasets/rasyosef/amharic-sentences-corpus.text1M<n<10M3 likes131 downloads2y agoHugging Face28gheero-Leyu /leyu-amharic-addis-ababa-dialect Leyu Amharic - Addis Ababa Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Addis Ababa dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-addis-ababa-dialect.audioautomatic-speech-recognition1K<n<10K1 likes127 downloads2mo agoHugging Face29gheero-Leyu /leyu-amharic-wello-dialect Leyu Amharic - Wello Dialect Speech Corpus Dataset Description This dataset is a curated parallel speech corpus consisting of audio recordings paired with corresponding text transcripts, focused on the Wello dialect of the Amharic language. It is designed to support speech technology research across multiple tasks, including Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The corpus captures dialect-specific phonetic variations, accent patterns, and… See the full description on the dataset page: https://huggingface.co/datasets/gheero-Leyu/leyu-amharic-wello-dialect.audioautomatic-speech-recognition10K<n<100K1 likes127 downloads2mo agoHugging Face30KYAGABA /amharic-speech-dataset-110HRS-V21audio10K<n<100K0 likes115 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.