datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
global-piqa-parallel
Global PIQA Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems.
In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.GPT-Paraphrases
GPT-Paraphrases dataset
This dataset contains text passages and their paraphrases generated using the GPT-3 language model. The paraphrases are designed to be semantically equivalent to the original text, but with different wording and structure.
The dataset includes text formatted in JSON and is in English.
Dataset Statistics
Number of text passages: Not specified in the information you provided.
Source of text passages: Not specified in the information you provided.… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/GPT-Paraphrases.
