datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
igbo-kongo_sentence-pairs
Igbo-Kongo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Kongo_Sentence-Pairs
Number of Rows: 68066
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-kongo_sentence-pairs.Igbo_Proverbs_Sample_V1
Tone-Marked Igbo Proverbs Parallel Corpus V1.1
50 manually curated Igbo proverbs with tone marks, dialect, context and zone metadata. This dataset will be regularly updated as we expand.
Useful for NLP, translation, and cultural preservation projects.
Collaboration, Licensing & Custom Services
This dataset is released under CC-BY-4.0 to support open research in Igbo NLP, MT, and cultural preservation. This 50-row proof-of-concept is, and will remain fully… See the full description on the dataset page: https://huggingface.co/datasets/Anyibaba/Igbo_Proverbs_Sample_V1.igbo_sa
Sentiment Analysis Data for the Igbo Language
Dataset Description:
This dataset contains a sentiment analysis dataset from Muhammad et al. (2023).
Data Structure:
The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages.
Citation:
@inproceedings{Muhammad2023AfriSentiAT,
title={AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages},
author={Shamsuddeen Hassan Muhammad and Idris Abdulmumin and Abinew Ali Ayele… See the full description on the dataset page: https://huggingface.co/datasets/DGurgurov/igbo_sa.igbo-chichewa_sentence-pairs
Igbo-Chichewa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Chichewa_Sentence-Pairs
Number of Rows: 529824
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-chichewa_sentence-pairs.bemba-igbo_sentence-pairs
Bemba-Igbo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bemba-Igbo_Sentence-Pairs
Number of Rows: 89452
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-igbo_sentence-pairs.igbo-pedi_sentence-pairs
Igbo-Pedi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Pedi_Sentence-Pairs
Number of Rows: 102943
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-pedi_sentence-pairs.igbo-tumbuka_sentence-pairs
Igbo-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Tumbuka_Sentence-Pairs
Number of Rows: 133589
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-tumbuka_sentence-pairs.igbo-twi_sentence-pairs
Igbo-Twi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Twi_Sentence-Pairs
Number of Rows: 127455
Number of Columns: 3… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-twi_sentence-pairs.igbo-ganda_sentence-pairs
Igbo-Ganda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Ganda_Sentence-Pairs
Number of Rows: 111322
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-ganda_sentence-pairs.fon-igbo_sentence-pairs
Fon-Igbo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Igbo_Sentence-Pairs
Number of Rows: 55155
Number of Columns: 3… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-igbo_sentence-pairs.ewe-igbo_sentence-pairs
Ewe-Igbo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ewe-Igbo_Sentence-Pairs
Number of Rows: 117699
Number of Columns: 3… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ewe-igbo_sentence-pairs.igbo-kikuyu_sentence-pairs
Igbo-Kikuyu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Kikuyu_Sentence-Pairs
Number of Rows: 55141
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-kikuyu_sentence-pairs.igbo-yoruba_sentence-pairs
Igbo-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Yoruba_Sentence-Pairs
Number of Rows: 414647
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-yoruba_sentence-pairs.igbo-tswana_sentence-pairs
Igbo-Tswana_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Tswana_Sentence-Pairs
Number of Rows: 159644
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-tswana_sentence-pairs.igbo-shona_sentence-pairs
Igbo-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Shona_Sentence-Pairs
Number of Rows: 498036
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-shona_sentence-pairs.igbo-kimbundu_sentence-pairs
Igbo-Kimbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Kimbundu_Sentence-Pairs
Number of Rows: 39828
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-kimbundu_sentence-pairs.bambara-igbo_sentence-pairs
Bambara-Igbo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bambara-Igbo_Sentence-Pairs
Number of Rows: 37303
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-igbo_sentence-pairs.igbo-swati_sentence-pairs
Igbo-Swati_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Swati_Sentence-Pairs
Number of Rows: 49075
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-swati_sentence-pairs.igbo-nuer_sentence-pairs
Igbo-Nuer_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Nuer_Sentence-Pairs
Number of Rows: 24786
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-nuer_sentence-pairs.hausa-igbo_sentence-pairs
Hausa-Igbo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Igbo_Sentence-Pairs
Number of Rows: 713539
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-igbo_sentence-pairs.IgboSenti-BBC
Dataset Description
This is a human-annotated sentiment dataset sourced from BBC for the Igbo language.
The dataset can be used for sentiment analysis tasks in Igbo languages.
How to Use
from datasets import load_dataset
dataset = load_dataset("Ifyokoh/IgboSenti-BBC")
igbo-tsonga_sentence-pairs
Igbo-Tsonga_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Tsonga_Sentence-Pairs
Number of Rows: 131392
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-tsonga_sentence-pairs.igbo-oromo_sentence-pairs
Igbo-Oromo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Oromo_Sentence-Pairs
Number of Rows: 63492
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-oromo_sentence-pairs.igbo-kinyarwanda_sentence-pairs
Igbo-Kinyarwanda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Kinyarwanda_Sentence-Pairs
Number of Rows: 181304… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-kinyarwanda_sentence-pairs.akan-igbo_sentence-pairs
Akan-Igbo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Akan-Igbo_Sentence-Pairs
Number of Rows: 39249
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-igbo_sentence-pairs.igbo-lingala_sentence-pairs
Igbo-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Lingala_Sentence-Pairs
Number of Rows: 93766
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-lingala_sentence-pairs.fulah-igbo_sentence-pairs
Fulah-Igbo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fulah-Igbo_Sentence-Pairs
Number of Rows: 111376
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fulah-igbo_sentence-pairs.igbo-xhosa_sentence-pairs
Igbo-Xhosa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Xhosa_Sentence-Pairs
Number of Rows: 375424
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-xhosa_sentence-pairs.igbo-tigrinya_sentence-pairs
Igbo-Tigrinya_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Tigrinya_Sentence-Pairs
Number of Rows: 101632
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-tigrinya_sentence-pairs.igbo-umbundu_sentence-pairs
Igbo-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Umbundu_Sentence-Pairs
Number of Rows: 58532
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-umbundu_sentence-pairs.
