datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rundi-tumbuka_sentence-pairs
Rundi-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Rundi-Tumbuka_Sentence-Pairs
Number of Rows: 194527
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/rundi-tumbuka_sentence-pairs.tigrinya-tumbuka_sentence-pairs
Tigrinya-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tigrinya-Tumbuka_Sentence-Pairs
Number of Rows: 152916… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tigrinya-tumbuka_sentence-pairs.igbo-tumbuka_sentence-pairs
Igbo-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Tumbuka_Sentence-Pairs
Number of Rows: 133589
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-tumbuka_sentence-pairs.fon-tumbuka_sentence-pairs
Fon-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Tumbuka_Sentence-Pairs
Number of Rows: 73794
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-tumbuka_sentence-pairs.swahili-tumbuka_sentence-pairs
Swahili-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Swahili-Tumbuka_Sentence-Pairs
Number of Rows: 499551
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-tumbuka_sentence-pairs.somali-tumbuka_sentence-pairs
Somali-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Tumbuka_Sentence-Pairs
Number of Rows: 179589
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-tumbuka_sentence-pairs.kamba-tumbuka_sentence-pairs
Kamba-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kamba-Tumbuka_Sentence-Pairs
Number of Rows: 63077
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kamba-tumbuka_sentence-pairs.dinka-tumbuka_sentence-pairs
Dinka-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dinka-Tumbuka_Sentence-Pairs
Number of Rows: 23524
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dinka-tumbuka_sentence-pairs.tumbuka-umbundu_sentence-pairs
Tumbuka-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tumbuka-Umbundu_Sentence-Pairs
Number of Rows: 99125
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tumbuka-umbundu_sentence-pairs.tumbuka-twi_sentence-pairs
Tumbuka-Twi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tumbuka-Twi_Sentence-Pairs
Number of Rows: 176324
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tumbuka-twi_sentence-pairs.pedi-tumbuka_sentence-pairs
Pedi-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Pedi-Tumbuka_Sentence-Pairs
Number of Rows: 101945
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/pedi-tumbuka_sentence-pairs.nuer-tumbuka_sentence-pairs
Nuer-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Nuer-Tumbuka_Sentence-Pairs
Number of Rows: 19926
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/nuer-tumbuka_sentence-pairs.kikuyu-tumbuka_sentence-pairs
Kikuyu-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kikuyu-Tumbuka_Sentence-Pairs
Number of Rows: 54854
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kikuyu-tumbuka_sentence-pairs.oromo-tumbuka_sentence-pairs
Oromo-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Tumbuka_Sentence-Pairs
Number of Rows: 73849
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-tumbuka_sentence-pairs.tsonga-tumbuka_sentence-pairs
Tsonga-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tsonga-Tumbuka_Sentence-Pairs
Number of Rows: 203106
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tsonga-tumbuka_sentence-pairs.shona-tumbuka_sentence-pairs
Shona-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Tumbuka_Sentence-Pairs
Number of Rows: 294060
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-tumbuka_sentence-pairs.hausa-tumbuka_sentence-pairs
Hausa-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Tumbuka_Sentence-Pairs
Number of Rows: 260765
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-tumbuka_sentence-pairs.fulah-tumbuka_sentence-pairs
Fulah-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fulah-Tumbuka_Sentence-Pairs
Number of Rows: 60296
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fulah-tumbuka_sentence-pairs.tswana-tumbuka_sentence-pairs
Tswana-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tswana-Tumbuka_Sentence-Pairs
Number of Rows: 187262
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tswana-tumbuka_sentence-pairs.bemba-tumbuka_sentence-pairs
Bemba-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bemba-Tumbuka_Sentence-Pairs
Number of Rows: 167288
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-tumbuka_sentence-pairs.swati-tumbuka_sentence-pairs
Swati-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Swati-Tumbuka_Sentence-Pairs
Number of Rows: 52465
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swati-tumbuka_sentence-pairs.kongo-tumbuka_sentence-pairs
Kongo-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kongo-Tumbuka_Sentence-Pairs
Number of Rows: 93758
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kongo-tumbuka_sentence-pairs.dyula-tumbuka_sentence-pairs
Dyula-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Tumbuka_Sentence-Pairs
Number of Rows: 69174
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-tumbuka_sentence-pairs.akan-tumbuka_sentence-pairs
Akan-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Akan-Tumbuka_Sentence-Pairs
Number of Rows: 36800
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-tumbuka_sentence-pairs.kimbundu-tumbuka_sentence-pairs
Kimbundu-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kimbundu-Tumbuka_Sentence-Pairs
Number of Rows: 63791… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kimbundu-tumbuka_sentence-pairs.ganda-tumbuka_sentence-pairs
Ganda-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ganda-Tumbuka_Sentence-Pairs
Number of Rows: 143005
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ganda-tumbuka_sentence-pairs.ewe-tumbuka_sentence-pairs
Ewe-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ewe-Tumbuka_Sentence-Pairs
Number of Rows: 178780
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ewe-tumbuka_sentence-pairs.chichewa-tumbuka_sentence-pairs
Chichewa-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Chichewa-Tumbuka_Sentence-Pairs
Number of Rows: 292147… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/chichewa-tumbuka_sentence-pairs.amharic-tumbuka_sentence-pairs
Amharic-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Amharic-Tumbuka_Sentence-Pairs
Number of Rows: 256777
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-tumbuka_sentence-pairs.lingala-tumbuka_sentence-pairs
Lingala-Tumbuka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Lingala-Tumbuka_Sentence-Pairs
Number of Rows: 129185
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-tumbuka_sentence-pairs.
