datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
akan-umbundu_sentence-pairs
Akan-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Akan-Umbundu_Sentence-Pairs
Number of Rows: 22651
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-umbundu_sentence-pairs.swahili-umbundu_sentence-pairs
Swahili-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Swahili-Umbundu_Sentence-Pairs
Number of Rows: 276674
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-umbundu_sentence-pairs.kongo-umbundu_sentence-pairs
Kongo-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kongo-Umbundu_Sentence-Pairs
Number of Rows: 58711
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kongo-umbundu_sentence-pairs.tsonga-umbundu_sentence-pairs
Tsonga-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tsonga-Umbundu_Sentence-Pairs
Number of Rows: 132324
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tsonga-umbundu_sentence-pairs.shona-umbundu_sentence-pairs
Shona-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Umbundu_Sentence-Pairs
Number of Rows: 153467
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-umbundu_sentence-pairs.oromo-umbundu_sentence-pairs
Oromo-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Umbundu_Sentence-Pairs
Number of Rows: 48783
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-umbundu_sentence-pairs.nuer-umbundu_sentence-pairs
Nuer-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Nuer-Umbundu_Sentence-Pairs
Number of Rows: 9607
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/nuer-umbundu_sentence-pairs.tumbuka-umbundu_sentence-pairs
Tumbuka-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tumbuka-Umbundu_Sentence-Pairs
Number of Rows: 99125
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tumbuka-umbundu_sentence-pairs.somali-umbundu_sentence-pairs
Somali-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Umbundu_Sentence-Pairs
Number of Rows: 96450
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-umbundu_sentence-pairs.kinyarwanda-umbundu_sentence-pairs
Kinyarwanda-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Umbundu_Sentence-Pairs
Number of Rows:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-umbundu_sentence-pairs.dinka-umbundu_sentence-pairs
Dinka-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dinka-Umbundu_Sentence-Pairs
Number of Rows: 11337
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dinka-umbundu_sentence-pairs.lingala-umbundu_sentence-pairs
Lingala-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Lingala-Umbundu_Sentence-Pairs
Number of Rows: 86575
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-umbundu_sentence-pairs.kikuyu-umbundu_sentence-pairs
Kikuyu-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kikuyu-Umbundu_Sentence-Pairs
Number of Rows: 29453
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kikuyu-umbundu_sentence-pairs.ganda-umbundu_sentence-pairs
Ganda-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ganda-Umbundu_Sentence-Pairs
Number of Rows: 76417
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ganda-umbundu_sentence-pairs.tswana-umbundu_sentence-pairs
Tswana-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tswana-Umbundu_Sentence-Pairs
Number of Rows: 109541
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tswana-umbundu_sentence-pairs.kimbundu-umbundu_sentence-pairs
Kimbundu-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kimbundu-Umbundu_Sentence-Pairs
Number of Rows: 47912… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kimbundu-umbundu_sentence-pairs.kamba-umbundu_sentence-pairs
Kamba-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kamba-Umbundu_Sentence-Pairs
Number of Rows: 41134
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kamba-umbundu_sentence-pairs.fon-umbundu_sentence-pairs
Fon-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Umbundu_Sentence-Pairs
Number of Rows: 56632
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-umbundu_sentence-pairs.ewe-umbundu_sentence-pairs
Ewe-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ewe-Umbundu_Sentence-Pairs
Number of Rows: 106645
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ewe-umbundu_sentence-pairs.bambara-umbundu_sentence-pairs
Bambara-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bambara-Umbundu_Sentence-Pairs
Number of Rows: 20038
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-umbundu_sentence-pairs.tigrinya-umbundu_sentence-pairs
Tigrinya-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tigrinya-Umbundu_Sentence-Pairs
Number of Rows: 66683… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tigrinya-umbundu_sentence-pairs.igbo-umbundu_sentence-pairs
Igbo-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Umbundu_Sentence-Pairs
Number of Rows: 58532
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-umbundu_sentence-pairs.fulah-umbundu_sentence-pairs
Fulah-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fulah-Umbundu_Sentence-Pairs
Number of Rows: 30013
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fulah-umbundu_sentence-pairs.dyula-umbundu_sentence-pairs
Dyula-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Umbundu_Sentence-Pairs
Number of Rows: 37912
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-umbundu_sentence-pairs.chichewa-umbundu_sentence-pairs
Chichewa-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Chichewa-Umbundu_Sentence-Pairs
Number of Rows: 141904… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/chichewa-umbundu_sentence-pairs.afrikaans-umbundu_sentence-pairs
Afrikaans-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Umbundu_Sentence-Pairs
Number of Rows: 205249… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-umbundu_sentence-pairs.rundi-umbundu_sentence-pairs
Rundi-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Rundi-Umbundu_Sentence-Pairs
Number of Rows: 121294
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/rundi-umbundu_sentence-pairs.bemba-umbundu_sentence-pairs
Bemba-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bemba-Umbundu_Sentence-Pairs
Number of Rows: 115464
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-umbundu_sentence-pairs.amharic-umbundu_sentence-pairs
Amharic-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Amharic-Umbundu_Sentence-Pairs
Number of Rows: 102087
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-umbundu_sentence-pairs.swati-umbundu_sentence-pairs
Swati-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Swati-Umbundu_Sentence-Pairs
Number of Rows: 32317
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swati-umbundu_sentence-pairs.
