datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tibetan_monolingual_A_filteredkinyarwanda_monolingual_v01.0
!!! PLEASE USE mbazaNLP/kinyarwanda_monolingual_v01.1 !!!
!!! This version contains several duplicates and few non-kinyarwanda documents
Dataset Summary
The Kinyarwanda Monolingual Dataset version 1 is a large collection of Kinyarwanda language texts aimed at supporting the development of NLP and AI applications which can process Kinyarwanda texts. This dataset contains 78k documents, totalling about 25 million words, and includes diverse content types such… See the full description on the dataset page: https://huggingface.co/datasets/mbazaNLP/kinyarwanda_monolingual_v01.0.tibetan_monolingual_A_metamari-monolingual-corpusA monolingual corpus of the Mari language in various genres, containing over 20 million word occurrences.
The presented genres:
Genre
Russian
English
мутер
словарь
dictionary
газетысе увер
газетные новости
periodical news
прозо
проза
prose
фольклор
фольклор
folklore
публицистике
публицистика
publicistic literature
поэзий
поэзия
poetry
трагикомедийтрагикомедия
tragicomedy
пьесе
пьеса
play
драме
драма
drama
комедий-водевиль
водевиль
vaudeville
комедий
комедия… See the full description on the dataset page: https://huggingface.co/datasets/mari-lab/mari-monolingual-corpus.grow-1-monolingual-1m-ha-en-scoredEmakhuwa-MonolingualBibTeX:
The dataset paper was published in EMNLP 2024.
Please cite as:
@inproceedings{ali-etal-2024-building,
title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks",
author = "Ali, Felermino D. M. A. and
Lopes Cardoso, Henrique and
Sousa-Silva, Rui",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Monolingual.
