datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
global-piqa-nonparallel
Global PIQA Non-Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The non-parallel split covers 136 language varieties, covering five continents, 18 language families, and 24 writing systems.
In this non-parallel split, over 50% of examples reference local foods, customs, traditions, or other culturally-specific elements.
Details are in our preprint:… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-nonparallel.global-piqa-parallel
Global PIQA Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems.
In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.nadora-global-industries
NADORA Global Industries
A synthetic multinational, built to be developed against rather than
demonstrated with.
One fictional company — $5.20bn revenue, $716m EBITDA, 24,000 employees, 18
countries, 35 legal entities, five business units — traded daily from January
2022 to December 2026 and rendered at six fidelities, from a 3 MB unit-test
fixture to a 5 GB full-scale corpus.
38,964,663 rows · 11 GB · 2,319 verification assertions, all passing.
100% synthetic. No real company… See the full description on the dataset page: https://huggingface.co/datasets/hemanthreddy901/nadora-global-industries.Global-LLMs-Replies
Global LLMs Replies
GPT-4o
-> 74,644 rows
mixtral-8x22b
-> 13,129 rows
claude-3-haiku
-> 3,871 rows
remote_sensing_VQA_multilingual
Remote Sensing VQA — Multilingual
A multilingual counterfactual MCQ dataset built from remote sensing / satellite imagery.
Each row contains a satellite image, two captions (original vs counterfactual), and a multiple-choice question probing whether a VLM follows the image or the misleading text.
Languages
Language
Code
Rows
English
en
50
Hindi
hi
50
Urdu
ur
50
Telugu
te
50
Bahasa Indonesia
id
50
Columns
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/apart-global-south-hack/remote_sensing_VQA_multilingual.global-censorship-index
Voidly Global Censorship Index
Real-time internet censorship measurements for 200 countries, based on 38,780,449+ OONI network probes.
Dataset Description
The Global Censorship Index provides country-level internet censorship scores derived from actual network measurements. Unlike annual expert assessments, this data updates daily.
Key Statistics
Countries covered: 200
Total measurements: 38,780,449
Severe censorship: 1 countries
High censorship: 6… See the full description on the dataset page: https://huggingface.co/datasets/emperor-mew/global-censorship-index.powertron-global-permafrost-corpus
Dataset Card: Powertron Global PermaFrost Corpus
Important Disambiguation: This corpus documents PermaFrost® NMR, a trademarked HVAC efficiency treatment product. It contains HVAC/refrigeration efficiency data (chillers, RTUs, DX systems, refrigeration). This corpus has NO connection to geological permafrost (frozen ground), climate science, or Arctic research. The name "PermaFrost" is a product trademark reflecting thermal transfer properties, not a geological term.… See the full description on the dataset page: https://huggingface.co/datasets/powertronglobal/powertron-global-permafrost-corpus.aya-global-exams-catalanCatalan exams for the Aya Global Exams.
Original data and file available here: link
Github Repo: link
counterfactual-pendulum-multilingual
📌 Dataset Summary
When a Vision-Language Model (VLM) is given an image along with a text prompt containing contradictory or misleading information, how does it react? Does it rely on the visual evidence, succumb to textual bias, or honestly abstain when faced with unresolvable conflict?
This dataset adapts the Counterfactual Pendulum scenario across two visual conflict dimensions:
Angular (Angle): Conflict in the pendulum's angle of inclination.
Light: Conflict in the light… See the full description on the dataset page: https://huggingface.co/datasets/apart-global-south-hack/counterfactual-pendulum-multilingual.prospire-synth-global-personas
🌍 Prospire Synth Global Personas
The World's Largest Unified Synthetic Persona Database
512M+ records · 82 columns · 77+ countries · 39 languages · DuckDB-native
🎯 What Is This?
Prospire Synth Global Personas is a unified, query-ready database of synthetic human personas built for AI agent simulations, market research, and cultural analysis. It merges 18 open-source datasets into a single coherent Parquet warehouse — partitioned, compressed, and… See the full description on the dataset page: https://huggingface.co/datasets/Kasher13/prospire-synth-global-personas.global-seo-knowledgeGlobal-Ocean-Science-Corpus
🌊 Global-Ocean-Science-Corpus (v2.0 Curated & Cleaned)
A Highly Curated, Large-Scale Pre-Training & RAG Corpus for Deep Ocean Sciences, Marine Biology, and Oceanography
Language Note: This dataset is a 100% English-language scientific corpus (language: "en") aggregating peer-reviewed literature, deep-sea exploration dossiers, and technical oceanographic reports from leading global marine institutes.
Global-Ocean-Science-Corpus, derin okyanus bilimleri, deniz biyolojisi… See the full description on the dataset page: https://huggingface.co/datasets/tilikumotp/Global-Ocean-Science-Corpus.actuarial-global-glossary-multilingual
🤝 Connect with me on LinkedIn!
Join the mission to make actuarial knowledge accessible worldwide
Let's discuss how AI can transform professional education and break language barriers in finance!
🌍 Global Actuarial Glossary - Breaking Language Barriers in Finance
🚀 The World's Most Comprehensive Multilingual Actuarial Dataset
Imagine: A brilliant actuarial student in Tokyo, a risk analyst in São Paulo, and an insurance executive… See the full description on the dataset page: https://huggingface.co/datasets/manuelcaccone/actuarial-global-glossary-multilingual.GlobalMedQA
GlobalMedQA — A Standardized Multilingual Dataset for Assessing Medical Knowledge in LLMs
Dataset Summary
GlobalMedQA is a harmonized multilingual dataset of medical multiple-choice questions (MCQs) designed to benchmark large language models in the healthcare domain.It integrates exam questions from 14 countries and 13 languages, standardized into a unified schema with consistent metadata and specialty classification based on the European Union of Medical Specialists… See the full description on the dataset page: https://huggingface.co/datasets/mariocedo/GlobalMedQA.Global_Environment-Social-And-Governance-Data
Global_Environment-Social-And-Governance Dataset
This Dataset contains all verified and authorized Environment, Social and Governance Statistics data in the World
Description
I have collected all data from WORLD-Bank's Data Catalog and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
https://datacatalog.worldbank.org/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Global_Environment-Social-And-Governance-Data.histoire-general-afrique-global-adaption
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
Svngoku/Histoire-General-Afrique-Global
This dataset contains French-language text excerpts detailing the political, social, and economic history of Africa from the 16th to the 18th centuries. The content covers specific regions such as the Lower Guinea Coast and the Zambezi, discussing topics like ethnic migrations, kingdom formations, and trade dynamics. Each sample consists… See the full description on the dataset page: https://huggingface.co/datasets/MaatAI/histoire-general-afrique-global-adaption.Global-MMLU-Lite-sample
Global-MMLU-Lite — Sample Subset
A small, fixed-size subset of CohereLabs/Global-MMLU-Lite intended for fast smoke-testing of multilingual MMLU evaluation pipelines.
What this is
40 examples per language, 15 languages, 600 examples total.
Strictly non-overlapping windows across languages: language i takes rows [i*40, i*40+40) from the per-language pool of test followed by dev (test=400, dev=200, pool=600). Concretely, languages 0–9 fall entirely within the original test… See the full description on the dataset page: https://huggingface.co/datasets/cyankiwi/Global-MMLU-Lite-sample.Global-MMLU-Lite
Global MMLU-Lite — Human Translated
Global MMLU-Lite is a multilingual
evaluation benchmark for LLMs covering 18 languages. This dataset extends it with professional
human translations for three additional low-resource languages that are not in the original:
Chichewa (nya), Māori (mri), and Inuktitut (iku).
Released as part of the BYOL: Bring Your Own Language Into LLMs
project (paper).
What's New
The original Global MMLU-Lite by Cohere
covers 18 languages: Arabic… See the full description on the dataset page: https://huggingface.co/datasets/ai-for-good-lab/Global-MMLU-Lite.global_mmlu_lite_pt
🌎 Global-MMLU Lite (Portuguese)
A Focused Benchmark for Portuguese-Language Reasoning in Large Language Models
Global-MMLU Lite (Portuguese) is a curated subset of the Global-MMLU Lite benchmark designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models in Portuguese, providing a diverse and computationally efficient collection of translated and adapted QA samples across domains such as general knowledge, science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_pt.aya-global-exams-spanishSpanish exams for the Aya Global Exams.
Original data and file available here: link
Github Repo: link
NOO-Verified-Global-Entities
🌐 NOO-Verified-Global-Entities (NOO AI互联网商业实体验证归属)
Official Data Infrastructure by Nooxus.com
[ 🚀 ANGEL ROUND INVESTOR NOTICE / 天使轮国际融资公告 ]
EN: Nooxus-AI is raising its Angel Round to scale nooxus.com — the world's first dedicated B2B Trading & Clearing Network for AI Agents. If your fund recognizes the trillion-dollar potential of building the "Visa / SWIFT network for the Agentic Web," powered by our production-ready 0.8ms RST signaling and Zero-Inbound stealth architecture… See the full description on the dataset page: https://huggingface.co/datasets/Nooxus-AI/NOO-Verified-Global-Entities.adaption-v5-global-employment-law-qa
WorkRight V5 — Global Employment Law QA
1,751 employment-law reasoning examples across five jurisdictions, with a
deterministic gold path and no language model anywhere in it.
sha256: 2e828767c42a9f5ff53083509bacf6b1313e86027d418048376ac9441e25a456
— byte-identical to the artefact evaluated on Adaption.
1. Dataset Summary
Every statutory answer here is produced by an executable rule that reads a
fact scenario and returns a structured record — eligibility, amount… See the full description on the dataset page: https://huggingface.co/datasets/sahilmaniyar888/adaption-v5-global-employment-law-qa.histoire-general-afrique-global-adaption
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
Svngoku/Histoire-General-Afrique-Global
This dataset contains French-language text excerpts detailing the political, social, and economic history of Africa from the 16th to the 18th centuries. The content covers specific regions such as the Lower Guinea Coast and the Zambezi, discussing topics like ethnic migrations, kingdom formations, and trade dynamics. Each sample consists… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/histoire-general-afrique-global-adaption.Educational-Flashcards-for-Global-Learners
1. Educational-Flashcards-for-Global-Learners/README.md
Educational Flashcards Dataset
Overview
A comprehensive collection of 100 educational flashcards covering STEM, humanities, law, arts, and cultural topics. Curated with 70% Indian content, 25% European, and 5% other Asian perspectives to promote diverse knowledge representation.
Dataset Structure
{
"input": "Text description",
"output": {
"type": "flashcards",
"topic": "Subject name"… See the full description on the dataset page: https://huggingface.co/datasets/Srinivasmec26/Educational-Flashcards-for-Global-Learners.GlobalPIQA_gl
GlobalPIQA_gl
Related paper: Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures
Dataset Summary
GlobalPIQA_gl is the Galician subset of GlobalPIQA, a multilingual benchmark for evaluating physical commonsense reasoning across more than 100 languages and cultural contexts. It is intended as an evaluation resource for models that must choose the most plausible solution to a practical physical situation.
The dataset follows the… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/GlobalPIQA_gl.law-global
The Case-law, centralizing legal decisions for better use, a community Dataset.
The Case-law Dataset is a comprehensive collection of legal decisons from various countries, centralized in a common format. This dataset aims to improve the development of legal AI models by providing a standardized, easily accessible corpus of global legal documents.
Join us in our mission to make AI more accessible and understandable for the legal world, ensuring that the power of language… See the full description on the dataset page: https://huggingface.co/datasets/gitrelief/law-global.tiny-aya-global-blindspots
tiny-aya-global: Blind Spot Evaluation
Model: CohereLabs/tiny-aya-global — ~3B parameter multilingual conversational model
Dataset: kelvinyelyen/tiny-aya-global-blindspots
Ten targeted probes designed to surface specific failure mechanisms, not aggregate accuracy. Each probe was run once under greedy decoding (temperature=0.0), then re-run 5x under sampling (temperature=0.7) to check whether each failure is a stable pattern or a one-off. Result: 7 clear failures, 1 pass, 1… See the full description on the dataset page: https://huggingface.co/datasets/kelvinyelyen/tiny-aya-global-blindspots.Global_Health-Nutrition-And-Population-Statistics
Global_Health-Nutrition-And-Population-Statistics Dataset
This Dataset contains all verified and authorized Health, Nutrition and Population Statistics data in the World
Description
I have collected all data from WORLD-Bank's Data Catalog and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
https://datacatalog.worldbank.org/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Global_Health-Nutrition-And-Population-Statistics.aya-global-exams-basqueSpanish exams for the Aya Global Exams.
Original data and file available here: link
Github Repo: link
global-mmlu-lite
Global MMLU Lite – Galician & Urdu
Machine-translated Galician and Urdu subsets of the Global MMLU Lite benchmark.
Dataset Description
Global MMLU Lite is a culturally-aware, multilingual evaluation benchmark for large language models, covering multiple-choice questions across many academic subjects. This repository contains Galician (gl) and Urdu (ur) translations. This dataset was translated using Google Machine Translate.
Splits
Config… See the full description on the dataset page: https://huggingface.co/datasets/Owos/global-mmlu-lite.
