datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Kinyarwanda_Engligh_Multilingual_ASRThis dataset was created from Mozilla's Common Voice dataset for the purposes of Multilingual ASR on Kinyarwanda and English.
The dataset contains 3000 hours of multilingual training samples, 300 hours of validation samples and 200 of testing samples.
kinyarwanda-tts-dataset
Kinyarwanda dataset for text to speech model
Kinyarwanda dataset for text to speech model holds data for ai modelling of Kinyarwanda chatbots or other use cases.
common-voice-kinyarwanda-english-dataset
Kinyarwanda-English Commonvoice dataset
A compilation of Kinyarwanda-english dataset to be used to train multi-lingual ASR
Note: The audio dataset shall be added in the future
Kinyarwanda_English_parallel_dataset
Kinyarwanda-English parallel text
This dataset contains 55,000 Kinyarwanda-English sentence pairs, obtained by scraping web data from religious sources such as:
Bible
Quran
This dataset has not been curated only cleaned.
dinka-kinyarwanda_sentence-pairs
Dinka-Kinyarwanda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dinka-Kinyarwanda_Sentence-Pairs
Number of Rows: 39196… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dinka-kinyarwanda_sentence-pairs.kinyarwanda-yoruba_sentence-pairs
Kinyarwanda-Yoruba_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Yoruba_Sentence-Pairs
Number of Rows: 229031… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-yoruba_sentence-pairs.kinyarwanda-nuer_sentence-pairs
Kinyarwanda-Nuer_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Nuer_Sentence-Pairs
Number of Rows: 31429… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-nuer_sentence-pairs.akan-kinyarwanda_sentence-pairs
Akan-Kinyarwanda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Akan-Kinyarwanda_Sentence-Pairs
Number of Rows: 67145… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-kinyarwanda_sentence-pairs.kinyarwanda-zulu_sentence-pairs
Kinyarwanda-Zulu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Zulu_Sentence-Pairs
Number of Rows: 546208… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-zulu_sentence-pairs.kinyarwanda-tigrinya_sentence-pairs
Kinyarwanda-Tigrinya_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Tigrinya_Sentence-Pairs
Number of Rows:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-tigrinya_sentence-pairs.Roleplay-Kinyarwanda
RolePlay-Kinyarwanda
Roleplay-Kinyarwanda Dataset is a dataset for roleplaying in the Kinyarwanda language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, it can be found at… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Kinyarwanda.kinyarwanda-pedi_sentence-pairs
Kinyarwanda-Pedi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Pedi_Sentence-Pairs
Number of Rows: 215097… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-pedi_sentence-pairs.amharic-kinyarwanda_sentence-pairs
Amharic-Kinyarwanda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Amharic-Kinyarwanda_Sentence-Pairs
Number of Rows:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-kinyarwanda_sentence-pairs.kinyarwanda-umbundu_sentence-pairs
Kinyarwanda-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Umbundu_Sentence-Pairs
Number of Rows:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-umbundu_sentence-pairs.kinyarwanda-twi_sentence-pairs
Kinyarwanda-Twi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Twi_Sentence-Pairs
Number of Rows: 230016
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-twi_sentence-pairs.kikuyu-kinyarwanda_sentence-pairs
Kikuyu-Kinyarwanda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kikuyu-Kinyarwanda_Sentence-Pairs
Number of Rows: 84433… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kikuyu-kinyarwanda_sentence-pairs.kamba-kinyarwanda_sentence-pairs
Kamba-Kinyarwanda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kamba-Kinyarwanda_Sentence-Pairs
Number of Rows: 74997… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kamba-kinyarwanda_sentence-pairs.igbo-kinyarwanda_sentence-pairs
Igbo-Kinyarwanda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Kinyarwanda_Sentence-Pairs
Number of Rows: 181304… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-kinyarwanda_sentence-pairs.english_kinyarwanda55,000 sentence translations, created by translating sentences found in the BYU English Corpora
(COCA)
Translated using DeepTranslate's Google Translate API
kinyarwanda-tswana_sentence-pairs
Kinyarwanda-Tswana_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Tswana_Sentence-Pairs
Number of Rows: 282264… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-tswana_sentence-pairs.kinyarwanda-oromo_sentence-pairs
Kinyarwanda-Oromo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Oromo_Sentence-Pairs
Number of Rows: 113822… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-oromo_sentence-pairs.kinyarwanda-xhosa_sentence-pairs
Kinyarwanda-Xhosa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Xhosa_Sentence-Pairs
Number of Rows: 366852… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-xhosa_sentence-pairs.kinyarwanda-wolof_sentence-pairs
Kinyarwanda-Wolof_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Wolof_Sentence-Pairs
Number of Rows: 68002… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-wolof_sentence-pairs.kinyarwanda-somali_sentence-pairs
Kinyarwanda-Somali_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Somali_Sentence-Pairs
Number of Rows: 268329… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-somali_sentence-pairs.fulah-kinyarwanda_sentence-pairs
Fulah-Kinyarwanda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fulah-Kinyarwanda_Sentence-Pairs
Number of Rows: 220054… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fulah-kinyarwanda_sentence-pairs.ewe-kinyarwanda_sentence-pairs
Ewe-Kinyarwanda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Ewe-Kinyarwanda_Sentence-Pairs
Number of Rows: 210357
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/ewe-kinyarwanda_sentence-pairs.dyula-kinyarwanda_sentence-pairs
Dyula-Kinyarwanda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dyula-Kinyarwanda_Sentence-Pairs
Number of Rows: 102925… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dyula-kinyarwanda_sentence-pairs.bemba-kinyarwanda_sentence-pairs
Bemba-Kinyarwanda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bemba-Kinyarwanda_Sentence-Pairs
Number of Rows: 197886… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-kinyarwanda_sentence-pairs.kinyarwanda-lingala_sentence-pairs
Kinyarwanda-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kinyarwanda-Lingala_Sentence-Pairs
Number of Rows:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kinyarwanda-lingala_sentence-pairs.hausa-kinyarwanda_sentence-pairs
Hausa-Kinyarwanda_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Hausa-Kinyarwanda_Sentence-Pairs
Number of Rows: 340233… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/hausa-kinyarwanda_sentence-pairs.
