datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yoruba_sa
Sentiment Analysis Data for the Yoruba Language
Dataset Description:
This dataset contains a sentiment analysis dataset from Muhammad et al. (2023).
Data Structure:
The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages.
Citation:
@inproceedings{Muhammad2023AfriSentiAT,
title={AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages},
author={Shamsuddeen Hassan Muhammad and Idris Abdulmumin and Abinew Ali… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/yoruba_sa.piqa_yoruba_pidgin
Physical Commonsense Reasoning for Yorùbá and Nigerian Pidgin
Dataset Summary
This dataset was developed for the MRL 2025 Shared Task on Multilingual Physical Reasoning. For more details, see Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures.
It provides a test collection for evaluating physical commonsense reasoning, that is, a model's ability to understand how objects, actions, and outcomes relate in everyday scenarios.
The… See the full description on the dataset page: https://huggingface.co/datasets/taresco/piqa_yoruba_pidgin.yorubaadagemenyo_20k_a_multi_domain_english_yoruba_corpus_for_machine_translationyoruba-english-pairskinyarwanda-yoruba_sentence-pairs
Kinyarwanda-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Yoruba_Sentence-Pairs
Number of Rows: 229031… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-yoruba_sentence-pairs.lingala-yoruba_sentence-pairs
Lingala-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Lingala-Yoruba_Sentence-Pairs
Number of Rows: 146711
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-yoruba_sentence-pairs.kimbundu-yoruba_sentence-pairs
Kimbundu-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kimbundu-Yoruba_Sentence-Pairs
Number of Rows: 69194
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kimbundu-yoruba_sentence-pairs.Yoruba-diacritics-vs-non-diacriticsbemba-yoruba_sentence-pairs
Bemba-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bemba-Yoruba_Sentence-Pairs
Number of Rows: 123086
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-yoruba_sentence-pairs.amharic-yoruba_sentence-pairs
Amharic-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Amharic-Yoruba_Sentence-Pairs
Number of Rows: 422788
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-yoruba_sentence-pairs.yoruba-diacritic-restoration-dataset
Yorùbá Diacritic Restoration Dataset
Prompt-completion pairs for restoring correct diacritical marks (tone marks and underdots) in undiacriticized Yorùbá text.
Dataset Details
Source: Derived from MENYO-20k, a multi-domain English–Yorùbá parallel corpus (news, TED talks, book excerpts, ICT, proverbs; JW-sourced content excluded)
Raw dataset: 8,365 rows, annotated with tone_pattern, harmony_class/harmony_breakdown, and focus_tag columns
Adapted dataset: 4,968 rows… See the full description on the dataset page: https://huggingface.co/datasets/Natashadonoh/yoruba-diacritic-restoration-dataset.yoruba_newsclass_topictsonga-yoruba_sentence-pairs
Tsonga-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tsonga-Yoruba_Sentence-Pairs
Number of Rows: 187996
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tsonga-yoruba_sentence-pairs.somali-yoruba_sentence-pairs
Somali-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Yoruba_Sentence-Pairs
Number of Rows: 378911
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-yoruba_sentence-pairs.igbo-yoruba_sentence-pairs
Igbo-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Yoruba_Sentence-Pairs
Number of Rows: 414647
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-yoruba_sentence-pairs.tswana-yoruba_sentence-pairs
Tswana-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tswana-Yoruba_Sentence-Pairs
Number of Rows: 213028
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tswana-yoruba_sentence-pairs.tigrinya-yoruba_sentence-pairs
Tigrinya-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tigrinya-Yoruba_Sentence-Pairs
Number of Rows: 133381
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tigrinya-yoruba_sentence-pairs.swahili-yoruba_sentence-pairs
Swahili-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Swahili-Yoruba_Sentence-Pairs
Number of Rows: 582991
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-yoruba_sentence-pairs.kamba-yoruba_sentence-pairs
Kamba-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kamba-Yoruba_Sentence-Pairs
Number of Rows: 71526
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kamba-yoruba_sentence-pairs.rundi-yoruba_sentence-pairs
Rundi-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Rundi-Yoruba_Sentence-Pairs
Number of Rows: 174926
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/rundi-yoruba_sentence-pairs.fulah-yoruba_sentence-pairs
Fulah-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fulah-Yoruba_Sentence-Pairs
Number of Rows: 374873
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fulah-yoruba_sentence-pairs.bambara-yoruba_sentence-pairs
Bambara-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bambara-Yoruba_Sentence-Pairs
Number of Rows: 61950
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-yoruba_sentence-pairs.shona-yoruba_sentence-pairs
Shona-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Yoruba_Sentence-Pairs
Number of Rows: 541537
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-yoruba_sentence-pairs.pedi-yoruba_sentence-pairs
Pedi-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Pedi-Yoruba_Sentence-Pairs
Number of Rows: 169985
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/pedi-yoruba_sentence-pairs.oromo-yoruba_sentence-pairs
Oromo-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Yoruba_Sentence-Pairs
Number of Rows: 84299
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-yoruba_sentence-pairs.kongo-yoruba_sentence-pairs
Kongo-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kongo-Yoruba_Sentence-Pairs
Number of Rows: 126349
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kongo-yoruba_sentence-pairs.dyula-yoruba_sentence-pairs
Dyula-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Yoruba_Sentence-Pairs
Number of Rows: 102298
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-yoruba_sentence-pairs.chichewa-yoruba_sentence-pairs
Chichewa-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Chichewa-Yoruba_Sentence-Pairs
Number of Rows: 525621
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/chichewa-yoruba_sentence-pairs.hausa-yoruba_sentence-pairs
Hausa-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Yoruba_Sentence-Pairs
Number of Rows: 795566
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-yoruba_sentence-pairs.
