datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
global-piqa-nonparallel
Global PIQA Non-Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The non-parallel split covers 136 language varieties, covering five continents, 18 language families, and 24 writing systems.
In this non-parallel split, over 50% of examples reference local foods, customs, traditions, or other culturally-specific elements.
Details are in our preprint:… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-nonparallel.global-piqa-parallel
Global PIQA Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems.
In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.Global_Environment-Social-And-Governance-Data
Global_Environment-Social-And-Governance Dataset
This Dataset contains all verified and authorized Environment, Social and Governance Statistics data in the World
Description
I have collected all data from WORLD-Bank's Data Catalog and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
https://datacatalog.worldbank.org/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Global_Environment-Social-And-Governance-Data.GlobalPIQA_gl
GlobalPIQA_gl
Related paper: Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures
Dataset Summary
GlobalPIQA_gl is the Galician subset of GlobalPIQA, a multilingual benchmark for evaluating physical commonsense reasoning across more than 100 languages and cultural contexts. It is intended as an evaluation resource for models that must choose the most plausible solution to a practical physical situation.
The dataset follows the… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/GlobalPIQA_gl.Global_Health-Nutrition-And-Population-Statistics
Global_Health-Nutrition-And-Population-Statistics Dataset
This Dataset contains all verified and authorized Health, Nutrition and Population Statistics data in the World
Description
I have collected all data from WORLD-Bank's Data Catalog and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
https://datacatalog.worldbank.org/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Global_Health-Nutrition-And-Population-Statistics.global-mmlu-lite
Global MMLU Lite – Galician & Urdu
Machine-translated Galician and Urdu subsets of the Global MMLU Lite benchmark.
Dataset Description
Global MMLU Lite is a culturally-aware, multilingual evaluation benchmark for large language models, covering multiple-choice questions across many academic subjects. This repository contains Galician (gl) and Urdu (ur) translations. This dataset was translated using Google Machine Translate.
Splits
Config… See the full description on the dataset page: https://huggingface.co/datasets/Owos/global-mmlu-lite.
