datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hausa-shona_sentence-pairs
Hausa-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Shona_Sentence-Pairs
Number of Rows: 829704
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-shona_sentence-pairs.amharic-shona_sentence-pairs
Amharic-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Amharic-Shona_Sentence-Pairs
Number of Rows: 856822
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-shona_sentence-pairs.shona-umbundu_sentence-pairs
Shona-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Umbundu_Sentence-Pairs
Number of Rows: 153467
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-umbundu_sentence-pairs.oromo-shona_sentence-pairs
Oromo-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Shona_Sentence-Pairs
Number of Rows: 102932
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-shona_sentence-pairs.fulah-shona_sentence-pairs
Fulah-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fulah-Shona_Sentence-Pairs
Number of Rows: 122877
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fulah-shona_sentence-pairs.kikuyu-shona_sentence-pairs
Kikuyu-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kikuyu-Shona_Sentence-Pairs
Number of Rows: 87949
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kikuyu-shona_sentence-pairs.igbo-shona_sentence-pairs
Igbo-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Shona_Sentence-Pairs
Number of Rows: 498036
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-shona_sentence-pairs.shona-tigrinya_sentence-pairs
Shona-Tigrinya_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Tigrinya_Sentence-Pairs
Number of Rows: 226951
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-tigrinya_sentence-pairs.chichewa-shona_sentence-pairs
Chichewa-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Chichewa-Shona_Sentence-Pairs
Number of Rows: 977418
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/chichewa-shona_sentence-pairs.bambara-shona_sentence-pairs
Bambara-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bambara-Shona_Sentence-Pairs
Number of Rows: 57467
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-shona_sentence-pairs.kongo-shona_sentence-pairs
Kongo-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kongo-Shona_Sentence-Pairs
Number of Rows: 127296
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kongo-shona_sentence-pairs.kimbundu-shona_sentence-pairs
Kimbundu-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kimbundu-Shona_Sentence-Pairs
Number of Rows: 90586
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kimbundu-shona_sentence-pairs.shona-twi_sentence-pairs
Shona-Twi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Twi_Sentence-Pairs
Number of Rows: 272927
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-twi_sentence-pairs.shona-tumbuka_sentence-pairs
Shona-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Tumbuka_Sentence-Pairs
Number of Rows: 294060
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-tumbuka_sentence-pairs.shona-tswana_sentence-pairs
Shona-Tswana_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Tswana_Sentence-Pairs
Number of Rows: 299901
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-tswana_sentence-pairs.shona-swati_sentence-pairs
Shona-Swati_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Swati_Sentence-Pairs
Number of Rows: 86651
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-swati_sentence-pairs.pedi-shona_sentence-pairs
Pedi-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Pedi-Shona_Sentence-Pairs
Number of Rows: 169322
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/pedi-shona_sentence-pairs.lingala-shona_sentence-pairs
Lingala-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Lingala-Shona_Sentence-Pairs
Number of Rows: 203472
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-shona_sentence-pairs.kamba-shona_sentence-pairs
Kamba-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kamba-Shona_Sentence-Pairs
Number of Rows: 94694
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kamba-shona_sentence-pairs.ewe-shona_sentence-pairs
Ewe-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ewe-Shona_Sentence-Pairs
Number of Rows: 281408
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ewe-shona_sentence-pairs.bemba-shona_sentence-pairs
Bemba-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bemba-Shona_Sentence-Pairs
Number of Rows: 245336
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-shona_sentence-pairs.afrikaans-shona_sentence-pairs
Afrikaans-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Shona_Sentence-Pairs
Number of Rows: 1293879… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-shona_sentence-pairs.shona-xhosa_sentence-pairs
Shona-Xhosa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Xhosa_Sentence-Pairs
Number of Rows: 847016
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-xhosa_sentence-pairs.shona-tsonga_sentence-pairs
Shona-Tsonga_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Tsonga_Sentence-Pairs
Number of Rows: 351101
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-tsonga_sentence-pairs.shona-swahili_sentence-pairs
Shona-Swahili_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Swahili_Sentence-Pairs
Number of Rows: 1099183
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-swahili_sentence-pairs.rundi-shona_sentence-pairs
Rundi-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Rundi-Shona_Sentence-Pairs
Number of Rows: 334206
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/rundi-shona_sentence-pairs.shona-yoruba_sentence-pairs
Shona-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Yoruba_Sentence-Pairs
Number of Rows: 541537
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-yoruba_sentence-pairs.nuer-shona_sentence-pairs
Nuer-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Nuer-Shona_Sentence-Pairs
Number of Rows: 35601
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/nuer-shona_sentence-pairs.dyula-shona_sentence-pairs
Dyula-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Shona_Sentence-Pairs
Number of Rows: 111253
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-shona_sentence-pairs.akan-shona_sentence-pairs
Akan-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Akan-Shona_Sentence-Pairs
Number of Rows: 68092
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-shona_sentence-pairs.
