datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Roleplay-Afrikaans
RolePlay-Afrikaans
Roleplay-Afrikaans Dataset is a dataset for roleplaying in the Afrikaans language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, it can be found at this… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Afrikaans.afrikaans-tswana_sentence-pairs
Afrikaans-Tswana_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Tswana_Sentence-Pairs
Number of Rows: 779259… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-tswana_sentence-pairs.afrikaans-wolof_sentence-pairs
Afrikaans-Wolof_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Wolof_Sentence-Pairs
Number of Rows: 237049
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-wolof_sentence-pairs.afrikaans-somali_sentence-pairs
Afrikaans-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Somali_Sentence-Pairs
Number of Rows: 1432536… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-somali_sentence-pairs.afrikaans-fulah_sentence-pairs
Afrikaans-Fulah_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Fulah_Sentence-Pairs
Number of Rows: 168995
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-fulah_sentence-pairs.afrikaans-dinka_sentence-pairs
Afrikaans-Dinka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Dinka_Sentence-Pairs
Number of Rows: 113795
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-dinka_sentence-pairs.afrikaans-tigrinya_sentence-pairs
Afrikaans-Tigrinya_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Tigrinya_Sentence-Pairs
Number of Rows: 454333… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-tigrinya_sentence-pairs.afrikaans-bambara_sentence-pairs
Afrikaans-Bambara_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Bambara_Sentence-Pairs
Number of Rows: 121709… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-bambara_sentence-pairs.afrikaans-ganda_sentence-pairs
Afrikaans-Ganda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Ganda_Sentence-Pairs
Number of Rows: 477046
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-ganda_sentence-pairs.afrikaans-akan_sentence-pairs
Afrikaans-Akan_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Akan_Sentence-Pairs
Number of Rows: 96786
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-akan_sentence-pairs.afrikaans-amharic_sentence-pairs
Afrikaans-Amharic_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Amharic_Sentence-Pairs
Number of Rows: 2084073… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-amharic_sentence-pairs.Glossaries_Sample_Afrikaans_SMCafrikaans-shona_sentence-pairs
Afrikaans-Shona_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Shona_Sentence-Pairs
Number of Rows: 1293879… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-shona_sentence-pairs.afrikaans-oromo_sentence-pairs
Afrikaans-Oromo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Oromo_Sentence-Pairs
Number of Rows: 471697
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-oromo_sentence-pairs.afrikaans-kikuyu_sentence-pairs
Afrikaans-Kikuyu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Kikuyu_Sentence-Pairs
Number of Rows: 127765… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-kikuyu_sentence-pairs.afrikaans-kamba_sentence-pairs
Afrikaans-Kamba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Kamba_Sentence-Pairs
Number of Rows: 99196
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-kamba_sentence-pairs.afrikaans-ewe_sentence-pairs
Afrikaans-Ewe_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Ewe_Sentence-Pairs
Number of Rows: 603870
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-ewe_sentence-pairs.afrikaans-swahili_sentence-pairs
Afrikaans-Swahili_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Swahili_Sentence-Pairs
Number of Rows: 2454179… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-swahili_sentence-pairs.afrikaans-nuer_sentence-pairs
Afrikaans-Nuer_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Nuer_Sentence-Pairs
Number of Rows: 51337
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-nuer_sentence-pairs.afrikaans-lingala_sentence-pairs
Afrikaans-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Lingala_Sentence-Pairs
Number of Rows: 346128… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-lingala_sentence-pairs.afrikaans-kongo_sentence-pairs
Afrikaans-Kongo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Kongo_Sentence-Pairs
Number of Rows: 199797
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-kongo_sentence-pairs.afrikaans-chichewa_sentence-pairs
Afrikaans-Chichewa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Chichewa_Sentence-Pairs
Number of Rows: 1149578… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-chichewa_sentence-pairs.afrikaans-yoruba_sentence-pairs
Afrikaans-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Yoruba_Sentence-Pairs
Number of Rows: 1775503… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-yoruba_sentence-pairs.afrikaans-xhosa_sentence-pairs
Afrikaans-Xhosa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Xhosa_Sentence-Pairs
Number of Rows: 1361577… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-xhosa_sentence-pairs.afrikaans-umbundu_sentence-pairs
Afrikaans-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Umbundu_Sentence-Pairs
Number of Rows: 205249… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-umbundu_sentence-pairs.afrikaans-tsonga_sentence-pairs
Afrikaans-Tsonga_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Tsonga_Sentence-Pairs
Number of Rows: 554519… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-tsonga_sentence-pairs.afrikaans-igbo_sentence-pairs
Afrikaans-Igbo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Igbo_Sentence-Pairs
Number of Rows: 820405
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-igbo_sentence-pairs.afrikaans-fon_sentence-pairs
Afrikaans-Fon_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Fon_Sentence-Pairs
Number of Rows: 250258
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-fon_sentence-pairs.afrikaans-dyula_sentence-pairs
Afrikaans-Dyula_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Dyula_Sentence-Pairs
Number of Rows: 130823
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-dyula_sentence-pairs.afrikaans-zulu_sentence-pairs
Afrikaans-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Zulu_Sentence-Pairs
Number of Rows: 1771310
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-zulu_sentence-pairs.
