datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kinyarwanda_monolingual_v01.1
task_categories:
- text-generation
language:
- rw
size_categories:
- 1K<n<10K
Dataset Summary
The Kinyarwanda Monolingual Dataset version 1 is a large collection of Kinyarwanda language texts aimed at supporting the development of NLP and AI applications which can process Kinyarwanda texts.
This dataset contains 1,068,161, with 63,001,765 words and includes diverse content types such as news articles, government reports, religious texts, legal documents, educational… See the full description on the dataset page: https://huggingface.co/datasets/mbazaNLP/kinyarwanda_monolingual_v01.1.kinyarwanda_monolingual_v01.0
!!! PLEASE USE mbazaNLP/kinyarwanda_monolingual_v01.1 !!!
!!! This version contains several duplicates and few non-kinyarwanda documents
Dataset Summary
The Kinyarwanda Monolingual Dataset version 1 is a large collection of Kinyarwanda language texts aimed at supporting the development of NLP and AI applications which can process Kinyarwanda texts. This dataset contains 78k documents, totalling about 25 million words, and includes diverse content types such… See the full description on the dataset page: https://huggingface.co/datasets/mbazaNLP/kinyarwanda_monolingual_v01.0.Code-170k-kinyarwanda
Dataset Description
Code-170k-kinyarwanda is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Kinyarwanda, making coding education accessible to Kinyarwanda speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Kinyarwanda language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-kinyarwanda.Roleplay-Kinyarwanda
RolePlay-Kinyarwanda
Roleplay-Kinyarwanda Dataset is a dataset for roleplaying in the Kinyarwanda language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, it can be found at… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Kinyarwanda.
