datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
afaan-oromoo-speech
Dataset.ET Afaan Oromoo Speech — v0.1.0
9.843 hours · 3,594 clips · 74 speakers · 3,283 distinct prompts
Dataset Summary
Read speech in Afaan Oromoo, crowdsourced from volunteer contributors in Ethiopia
through a Telegram bot, peer-validated by other contributors, and screened
acoustically before release. Afaan Oromoo has very little open speech data; this
corpus exists to change that.
Contributors read a displayed prompt aloud, other contributors listen and vote… See the full description on the dataset page: https://huggingface.co/datasets/snapwre/afaan-oromoo-speech.afaan_oromo-speechafaan-oromo-tts
Afaan Oromo TTS
Configs
default
Splits
train
Columns
id
speaker_id
text
transcription
language
gender
audio
amharic-oromo_sentence-pairs
Amharic-Oromo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Amharic-Oromo_Sentence-Pairs
Number of Rows: 109805
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-oromo_sentence-pairs.oromo_name
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Soressaa/oromo_name.english-oromo_sentence-pairs_mt560
English-Oromo Parallel Dataset
This dataset contains parallel sentences in English and Oromo (orm).
Dataset Information
Language Pair: English ↔ Oromo
Language Code: orm
Country: orm
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this dataset, please cite the citation guide of the… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-oromo_sentence-pairs_mt560.Code-170k-oromo
Dataset Description
Code-170k-oromo is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Oromo, making coding education accessible to Oromo speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Oromo language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-oromo.english-oromo_sentence-pairs
English-Oromo_Sentence-Pairs Dataset
This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks.
It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: English-Oromo_Sentence-Pairs
File Size: 373356060 bytes
Languages: English, English
Dataset Description
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/zarnite/english-oromo_sentence-pairs.oromo_cleaned_datasetafaan_oromo_dataset_finaloromo_audioalpaca_oromo_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_oromo_taco.pedi-oromo_sentence-pairs
Pedi-Oromo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Pedi-Oromo_Sentence-Pairs
Number of Rows: 80150
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/pedi-oromo_sentence-pairs.french-oromo_sentence-pairs
French-Oromo_Sentence-Pairs Dataset
This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks.
It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: French-Oromo_Sentence-Pairs
File Size: 113902124 bytes
Languages: French, French
Dataset Description
The dataset contains sentence pairs in… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/french-oromo_sentence-pairs.oromo-sentiments-corpus
Oromo Sentiment Corpus
Dataset Description
This dataset contains sentiment-labeled text data in Oromo for binary sentiment classification (Positive/Negative). Sentiments are extracted and processed from the English meanings of the sentences using DistilBERT for sentiment classification. The dataset is part of a larger collection of African language sentiment analysis resources.
Dataset Statistics
Total samples: 825,272
Positive sentiment: 457396 (55.4%)… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-sentiments-corpus.leyu-oromo-speech-corpus-2026
Leyu Afaan Oromo Speech Corpus 2026
Official speech dataset submission for the Leyu Data Collection Competition 2026.
Organization & Team
Hugging Face Org: SoundWaveET
Dataset Repo: SoundWaveET/leyu-oromo-speech-corpus-2026
oromo-xhosa_sentence-pairs
Oromo-Xhosa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Xhosa_Sentence-Pairs
Number of Rows: 104058
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-xhosa_sentence-pairs.oromo-somali_sentence-pairs
Oromo-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Somali_Sentence-Pairs
Number of Rows: 88097
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-somali_sentence-pairs.oromo-shona_sentence-pairs
Oromo-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Shona_Sentence-Pairs
Number of Rows: 102932
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-shona_sentence-pairs.oromo-emotions-corpus
Oromo Emotion Analysis Corpus
Dataset Description
This dataset contains emotion-labeled text data in Oromo for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources.
Dataset Statistics
Total samples: 825… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-emotions-corpus.english-oromo_sentence-pairs
English-Oromo_Sentence-Pairs Dataset
This dataset can be used for machine translation, sentence alignment, or other natural language processing tasks.
It is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: English-Oromo_Sentence-Pairs
File Size: 373356060 bytes
Languages: English, English
Dataset Description
The dataset contains sentence pairs in… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-oromo_sentence-pairs.alpaca-oromo-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-oromo-cleaned.oromo-wolof_sentence-pairs
Oromo-Wolof_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Wolof_Sentence-Pairs
Number of Rows: 38907
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-wolof_sentence-pairs.akan-oromo_sentence-pairs
Akan-Oromo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Akan-Oromo_Sentence-Pairs
Number of Rows: 47083
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-oromo_sentence-pairs.afaan_oromo_datasetoromo-rundi_sentence-pairs
Oromo-Rundi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Rundi_Sentence-Pairs
Number of Rows: 91889
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-rundi_sentence-pairs.kimbundu-oromo_sentence-pairs
Kimbundu-Oromo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kimbundu-Oromo_Sentence-Pairs
Number of Rows: 34479
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kimbundu-oromo_sentence-pairs.Oromo_datasetoromo-umbundu_sentence-pairs
Oromo-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Umbundu_Sentence-Pairs
Number of Rows: 48783
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-umbundu_sentence-pairs.fulah-oromo_sentence-pairs
Fulah-Oromo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fulah-Oromo_Sentence-Pairs
Number of Rows: 92794
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fulah-oromo_sentence-pairs.
