datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
global-piqa-parallel
Global PIQA Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems.
In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.GSM8KInstruct_ParallelWenYanWen_English_Parallel
Dataset Card for WenYanWen_English_Parallel
Dataset Summary
The WenYanWen_English_Parallel dataset is a multilingual parallel corpus in Classical Chinese (Wenyanwen), modern Chinese, and English. The Classical Chinese and modern Chinese parts are sourced from the NiuTrans/Classical-Modern dataset, while the corresponding English translations are generated using Gemini Pro.
Data Fields
info: A string representing the title or source information of the text.… See the full description on the dataset page: https://huggingface.co/datasets/KaifengGGG/WenYanWen_English_Parallel.mlqa_parallel_ar_eg
Dataset Card for "mlqa_parallel_ar_eg"
More Information needed
welsh_parallel_corpora
🏴🇬🇧 Welsh-English Parallel Corpora Translation Dataset
A curated bidirectional translation dataset containing 324,904 Welsh-English parallel sentences in chat format, designed for fine-tuning language models on low-resource language translation.
Please find a blog on the data curation process here.
Dataset Description
This dataset provides Welsh-English translation pairs from multiple parallel corpora sources. Welsh (Cymraeg) is a low-resource language… See the full description on the dataset page: https://huggingface.co/datasets/locailabs/welsh_parallel_corpora.
