CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yordanoswuletaw /amharic-pretraining-corpusAmharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining tasks. It consists of diverse text sources, including news articles, books, social media posts, government documents, and web content, all written in Amharic. You can load the dataset as follows from datasets import load_dataset ds = load_dataset("yordanoswuletaw/amharic-pretraining-corpus") texttext-generation100M<n<1B4 likes293 downloads2y agoHugging Face02uhhlt /amharichatespeechranlp Introduction The Amharic Hate Speech data is collected using the Twitter API spanning from October 1, 2020 - November 30, 2022, considering the socio-political dynamics of Ethiopia in Twitter space. We used WebAnno tool for data annotation; each tweet is annotated by two native speakers and curated by one more experienced adjudicator to determine the gold labels. A total of 15.1k tweets consisting of three class labels namely: Hate, Offensive and Normal are presented. Read our… See the full description on the dataset page: https://huggingface.co/datasets/uhhlt/amharichatespeechranlp.texttext-classification10K<n<100K1 likes199 downloads2y agoHugging Face03AddisGPT /AddisGPT-Amharic-Instruction AddisGPT-Amharic-Instruction A human-verified, fully conversational Amharic instruction-tuning dataset sourced entirely from real AddisGPT user interactions. 796 curated instruction–output pairs spanning 14 topics, drawn exclusively from anonymized conversations with AddisGPT — an Amharic-first AI assistant serving Ethiopian and diaspora communities. Every pair is an organic user question paired with the assistant's response; there is no synthetic, templated, or third-party… See the full description on the dataset page: https://huggingface.co/datasets/AddisGPT/AddisGPT-Amharic-Instruction.tabulartext-generationn<1K1 likes105 downloads25d agoHugging Face04israel /Amharic-News-Text-classification-Dataset An Amharic News Text classification Dataset In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.tabular10K<n<100K1 likes89 downloads4y agoHugging Face05Tvsybkzkmapab /Amharic_ad_generationtext1K<n<10K0 likes86 downloads3y agoHugging Face06michsethowusu /amharic-oromo_sentence-pairs Amharic-Oromo_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Amharic-Oromo_Sentence-Pairs Number of Rows: 109805 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-oromo_sentence-pairs.text100K<n<1M1 likes62 downloads1y agoHugging Face07mrxyz518 /Amharic_datasettabular100K<n<1M0 likes46 downloads29d agoHugging Face08ofc-its-phyla /amharic-speech-dataset-2026 Amharic Speech Dataset 2026 Overview This dataset contains Amharic speech recordings collected using the Leyu Platform for the Leyu Platform Competition 2026. Language Amharic (am) Dialect Standard Addis Ababa Amharic Speaker Information Number of Speakers: 1 Speaker IDs: SPK001 Audio Format Format: M4A Duration: 10–60 seconds per recording Directory Structure audio/ metadata.csv… See the full description on the dataset page: https://huggingface.co/datasets/ofc-its-phyla/amharic-speech-dataset-2026.audion<1K0 likes40 downloads2mo agoHugging Face09israel /AmharicStoryQAtabular1K<n<10K2 likes32 downloads1y agoHugging Face10rcade /amharic_wordstabular10K<n<100K1 likes29 downloads3y agoHugging Face11DGurgurov /amharic_sa Sentiment Analysis Data for the Amharic Language Dataset Description: This dataset contains a sentiment analysis dataset from Tesfa et al. (2024). Data Structure: The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages. Citation: @inproceedings{tesfa2024aspect, title={Aspect-Based Sentiment Analysis on Amharic Text for Evaluating Ethio-Telecom Services}, author={Tesfa, Tarikwa and Belete, Befikadu and Abera, Samuel and… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/amharic_sa.texttext-classification1K<n<10K0 likes29 downloads2y agoHugging Face12Henok /aya_amharic_dataset AYA Amharic Dataset This is the Amharic-only extract from the Aya Dataset, a multilingual instruction fine-tuning dataset. By @henok Dataset Summary The Aya Dataset is a multilingual instruction fine-tuning dataset curated by an open-science community via Aya Annotation Platform from Cohere For AI. The dataset contains a total of 204k human-annotated prompt-completion pairs along with the demographics data of the annotators. This dataset can be used to train, finetune… See the full description on the dataset page: https://huggingface.co/datasets/Henok/aya_amharic_dataset.text1K<n<10K0 likes26 downloads2y agoHugging Face13michsethowusu /amharic-shona_sentence-pairs Amharic-Shona_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Amharic-Shona_Sentence-Pairs Number of Rows: 856822 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-shona_sentence-pairs.text100K<n<1M0 likes25 downloads1y agoHugging Face14michsethowusu /amharic-ganda_sentence-pairs Amharic-Ganda_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Amharic-Ganda_Sentence-Pairs Number of Rows: 179444 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-ganda_sentence-pairs.text100K<n<1M0 likes21 downloads1y agoHugging Face15michsethowusu /amharic-fulah_sentence-pairs Amharic-Fulah_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Amharic-Fulah_Sentence-Pairs Number of Rows: 435048 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-fulah_sentence-pairs.text100K<n<1M0 likes21 downloads1y agoHugging Face16israel /AmharicZefentextn<1K0 likes20 downloads3y agoHugging Face17israel /AmharicSpellChecktext10K<n<100K0 likes18 downloads3y agoHugging Face18mHossain /Amharic_sum_finaltext10K<n<100K0 likes17 downloads3y agoHugging Face19michsethowusu /amharic-nuer_sentence-pairs Amharic-Nuer_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Amharic-Nuer_Sentence-Pairs Number of Rows: 35192 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-nuer_sentence-pairs.text10K<n<100K0 likes17 downloads1y agoHugging Face20michsethowusu /amharic-kimbundu_sentence-pairs Amharic-Kimbundu_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Amharic-Kimbundu_Sentence-Pairs Number of Rows: 64624… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-kimbundu_sentence-pairs.text10K<n<100K0 likes17 downloads1y agoHugging Face21michsethowusu /amharic-swati_sentence-pairs Amharic-Swati_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Amharic-Swati_Sentence-Pairs Number of Rows: 66047 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-swati_sentence-pairs.text10K<n<100K0 likes16 downloads1y agoHugging Face22michsethowusu /amharic-hausa_sentence-pairs Amharic-Hausa_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Amharic-Hausa_Sentence-Pairs Number of Rows: 751955 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-hausa_sentence-pairs.text100K<n<1M0 likes16 downloads1y agoHugging Face23israel /AmharicQAtext1K<n<10K0 likes15 downloads3y agoHugging Face24adunca08 /amharic_datatext1K<n<10K0 likes15 downloads2y agoHugging Face25michsethowusu /amharic-yoruba_sentence-pairs Amharic-Yoruba_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Amharic-Yoruba_Sentence-Pairs Number of Rows: 422788 Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-yoruba_sentence-pairs.text100K<n<1M0 likes15 downloads1y agoHugging Face26michsethowusu /amharic-xhosa_sentence-pairs Amharic-Xhosa_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Amharic-Xhosa_Sentence-Pairs Number of Rows: 520519 Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-xhosa_sentence-pairs.text100K<n<1M0 likes15 downloads1y agoHugging Face27michsethowusu /amharic-kinyarwanda_sentence-pairs Amharic-Kinyarwanda_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Amharic-Kinyarwanda_Sentence-Pairs Number of Rows:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-kinyarwanda_sentence-pairs.text100K<n<1M0 likes15 downloads1y agoHugging Face28michsethowusu /amharic-tswana_sentence-pairs Amharic-Tswana_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Amharic-Tswana_Sentence-Pairs Number of Rows: 278274 Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-tswana_sentence-pairs.text100K<n<1M0 likes12 downloads1y agoHugging Face29michsethowusu /afrikaans-amharic_sentence-pairs Afrikaans-Amharic_Sentence-Pairs Dataset This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks. This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1 Metadata File Name: Afrikaans-Amharic_Sentence-Pairs Number of Rows: 2084073… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-amharic_sentence-pairs.text1M<n<10M0 likes12 downloads1y agoHugging Face30dagn /expanded-amharic-news-datasetgated Expanded Amharic News Dataset (2011–2024) Dataset Description The Expanded Amharic News Dataset is a large-scale, ethically collected corpus of Amharic-language news articles written in Geʽez (Fidel) script, designed to support research in Natural Language Processing (NLP). This dataset builds upon the “An Amharic News Text Classification Dataset” developed by Israel Abebe Azime and Nebil Mohammed (arXiv link), which categorized Amharic news articles into multiple topical… See the full description on the dataset page: https://huggingface.co/datasets/dagn/expanded-amharic-news-dataset.texttext-classification100K<n<1M0 likes12 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.