datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NMT_Rwandan-Gazette_parallel_data_en_kin
Dataset Details
Dataset Description
This is a curated parallel dataset from the Official Gazette of the Republic of Rwanda. It has been curated to extract corresponding English and Kinyarwanda text and in the future we shall add French to the mix
Curated by: Digital Umuganda
Language(s) (NLP): Kinyarwanda and English
License: cc-by-4.0
Dataset Sources [optional]
The dataset original content was retrieved from the Rwandan ministry of Justice website… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/NMT_Rwandan-Gazette_parallel_data_en_kin.llama_nmt
중-한 번역
subset: ch-ko_basic_science
length: 37.7k
subset: ch-ko_broadcast
length: 362k
subset: ch-ko_daily_colloquial
length: 600k
subset: ch-ko_food
length: 1.2M
subset: ch-ko_humanities
length: 33.4k
subset: ch-ko_utterance_type
length: 12k
영-한 번역
subset: en-ko_basic_science
length: 356
subset: en-ko_broadcast
length: 121k
subset: en-ko_daily_colloquial
length: 1.2M
subset: en-ko_food
length: 1.2M
subset: en-ko_humanities
length:… See the full description on the dataset page: https://huggingface.co/datasets/wisenut-nlp-team/llama_nmt.NMT_Health_parallel_data_en_kinNMT_Tourism_parallel_data_en_kin
Dataset Description
This dataset was created in an effort to create a machine translation model for English-to-Kinyarwanda translation and vice-versa in a tourism-geared context.
Repository:link to the GitHub repository containing the code for training the model on this data, and the code for the collection of the monolingual data.
Data Format: TSV
Data Source: web scraping, manual annotation
Model: huggingface model link.
Data Instances
25375 49363 21210 Bird watching… See the full description on the dataset page: https://huggingface.co/datasets/mbazaNLP/NMT_Tourism_parallel_data_en_kin.
