datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nuer-zulu_sentence-pairs
Nuer-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Nuer-Zulu_Sentence-Pairs
Number of Rows: 38861
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/nuer-zulu_sentence-pairs.chichewa-zulu_sentence-pairs
Chichewa-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Chichewa-Zulu_Sentence-Pairs
Number of Rows: 1184038
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/chichewa-zulu_sentence-pairs.swahili-zulu_sentence-pairs
Swahili-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Swahili-Zulu_Sentence-Pairs
Number of Rows: 1366197
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-zulu_sentence-pairs.tsonga-zulu_sentence-pairs
Tsonga-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tsonga-Zulu_Sentence-Pairs
Number of Rows: 548608
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tsonga-zulu_sentence-pairs.swati-zulu_sentence-pairs
Swati-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Swati-Zulu_Sentence-Pairs
Number of Rows: 193267
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swati-zulu_sentence-pairs.somali-zulu_sentence-pairs
Somali-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Somali-Zulu_Sentence-Pairs
Number of Rows: 605556
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/somali-zulu_sentence-pairs.kongo-zulu_sentence-pairs
Kongo-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kongo-Zulu_Sentence-Pairs
Number of Rows: 155908
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kongo-zulu_sentence-pairs.kinyarwanda-zulu_sentence-pairs
Kinyarwanda-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Zulu_Sentence-Pairs
Number of Rows: 546208… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-zulu_sentence-pairs.zulu-pretraining-datasetThis is IsiZulu Pretraining Dataset. The dataset was used to pre-train BafoGPT-3B
Books: Zulu-English Dictionary – A dictionary offering Zulu terms with English definitions, ideal for teaching basic word mappings.
Translation: South African Government Speeches – Official speeches in Zulu, which help the model understand structured Zulu sentences and phrases.
Transcription: Zulu Community Corpus – A collection of transcriptions, exposing the model to real-life conversational Zulu.
Document:… See the full description on the dataset page: https://huggingface.co/datasets/ChallengerSpaceShuttle/zulu-pretraining-dataset.xhosa-zulu_sentence-pairs
Xhosa-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Xhosa-Zulu_Sentence-Pairs
Number of Rows: 1249452
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/xhosa-zulu_sentence-pairs.ewe-zulu_sentence-pairs
Ewe-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ewe-Zulu_Sentence-Pairs
Number of Rows: 344559
Number of Columns: 3… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ewe-zulu_sentence-pairs.alpaca-zuluoromo-zulu_sentence-pairs
Oromo-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Oromo-Zulu_Sentence-Pairs
Number of Rows: 120892
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/oromo-zulu_sentence-pairs.tswana-zulu_sentence-pairs
Tswana-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tswana-Zulu_Sentence-Pairs
Number of Rows: 412361
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tswana-zulu_sentence-pairs.akan-zulu_sentence-pairs
Akan-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Akan-Zulu_Sentence-Pairs
Number of Rows: 82123
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-zulu_sentence-pairs.pedi-zulu_sentence-pairs
Pedi-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Pedi-Zulu_Sentence-Pairs
Number of Rows: 313999
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/pedi-zulu_sentence-pairs.kikuyu-zulu_sentence-pairs
Kikuyu-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kikuyu-Zulu_Sentence-Pairs
Number of Rows: 106379
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kikuyu-zulu_sentence-pairs.hausa-zulu_sentence-pairs
Hausa-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Zulu_Sentence-Pairs
Number of Rows: 964410
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-zulu_sentence-pairs.Roleplay-Zulu
RolePlay-Zulu
Roleplay-Zulu Dataset is a dataset for roleplaying in the Zulu language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, it can be found at this github repo.
For… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Zulu.lingala-zulu_sentence-pairs
Lingala-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Lingala-Zulu_Sentence-Pairs
Number of Rows: 285795
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/lingala-zulu_sentence-pairs.fon-zulu_sentence-pairs
Fon-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fon-Zulu_Sentence-Pairs
Number of Rows: 137778
Number of Columns: 3… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fon-zulu_sentence-pairs.bemba-zulu_sentence-pairs
Bemba-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bemba-Zulu_Sentence-Pairs
Number of Rows: 312412
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-zulu_sentence-pairs.Glossaries_Sample_Zulu_SMCtigrinya-zulu_sentence-pairs
Tigrinya-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Tigrinya-Zulu_Sentence-Pairs
Number of Rows: 251069
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/tigrinya-zulu_sentence-pairs.kamba-zulu_sentence-pairs
Kamba-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kamba-Zulu_Sentence-Pairs
Number of Rows: 121895
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kamba-zulu_sentence-pairs.bambara-zulu_sentence-pairs
Bambara-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bambara-Zulu_Sentence-Pairs
Number of Rows: 66614
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bambara-zulu_sentence-pairs.afrikaans-zulu_sentence-pairs
Afrikaans-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Zulu_Sentence-Pairs
Number of Rows: 1771310
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-zulu_sentence-pairs.shona-zulu_sentence-pairs
Shona-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Shona-Zulu_Sentence-Pairs
Number of Rows: 1309315
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/shona-zulu_sentence-pairs.rundi-zulu_sentence-pairs
Rundi-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Rundi-Zulu_Sentence-Pairs
Number of Rows: 467223
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/rundi-zulu_sentence-pairs.kimbundu-zulu_sentence-pairs
Kimbundu-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kimbundu-Zulu_Sentence-Pairs
Number of Rows: 126319
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kimbundu-zulu_sentence-pairs.
