datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
monolingual_machine_translation_dataNMT_Rwandan-Gazette_parallel_data_en_kin
Dataset Details
Dataset Description
This is a curated parallel dataset from the Official Gazette of the Republic of Rwanda. It has been curated to extract corresponding English and Kinyarwanda text and in the future we shall add French to the mix
Curated by: Digital Umuganda
Language(s) (NLP): Kinyarwanda and English
License: cc-by-4.0
Dataset Sources [optional]
The dataset original content was retrieved from the Rwandan ministry of Justice website… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/NMT_Rwandan-Gazette_parallel_data_en_kin.Afrivoice_Swahili-Voice_Instruct_FormatASR_Fellowship_Challenge_Datasetkinyarwanda-tts-dataset
Kinyarwanda dataset for text to speech model
Kinyarwanda dataset for text to speech model holds data for ai modelling of Kinyarwanda chatbots or other use cases.
DigitalUmuganda_AfriVoice_shonacommon-voice-kinyarwanda-text-dataset
Dataset Card for DigitalUmuganda/common-voice-kinyarwanda-text-dataset
NMT_Health_parallel_data_en_kinMonolingual_health_dataset
Monolingual Dataset
This a a malnutrition dataset in Kinyarwanda and English, it shall be translated using translators to make it a parallel corpus.
Source of Data
Rwanda Biomedical Center (RBC) (26,390 sentences)
GPT-4 prompting (42,576 sentences)
Text_Normalization_Challenge_Unittests_Eng_FraAfrivoice_Swahili_0.0
Dataset summary
Domain
Total number of hours
Total number of transcribed hours
Total number of clips
Total Size of the dataset in GB
Agriculture
35.19
35.19
5,852
2.6
Health
62.11
62.11
10,315
6.4
Finance
118.75
118.75
19,629
16.6
Government
106.41
106.41
17,530
12.1
Education
91.96
91.96
15,204
6.2
Total
414.42
414.42
68,530
43.9
How to use
The datasets library allows you to load and pre-process your dataset in pure Python, at scale. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/Afrivoice_Swahili_0.0.kin_en_DigitalUmuganda
Dataset Card for "kin_en_DigitalUmuganda"
More Information needed
Dataset Information
The dataset was created by DigitalUmuganda for machine translation from Kinyarwanda to English
anv_test_data_nt_swahiliAfrivoice_Kinyarwanda_old_version
Dataset summary
[need more information]
Supported tasks
[need more information]
How to use
[need more information]
Dataset structure
Data fields
[need more information]
Data splits
[need more information]
Data preprocessing
[need more information]
Licensing Information
All datasets are licensed under the Creative Commons license (CC-BY-4).
