CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mrlbenchmarks /global-piqa-nonparallel Global PIQA Non-Parallel Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The non-parallel split covers 136 language varieties, covering five continents, 18 language families, and 24 writing systems. In this non-parallel split, over 50% of examples reference local foods, customs, traditions, or other culturally-specific elements. Details are in our preprint:… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-nonparallel.imagequestion-answering10K<n<100K40 likes5.9k downloads4mo agoHugging Face02mrlbenchmarks /global-piqa-parallel Global PIQA Parallel Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems. In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.imagequestion-answering10K<n<100K10 likes3.6k downloads4mo agoHugging Face03hemanthreddy901 /nadora-global-industries NADORA Global Industries A synthetic multinational, built to be developed against rather than demonstrated with. One fictional company — $5.20bn revenue, $716m EBITDA, 24,000 employees, 18 countries, 35 legal entities, five business units — traded daily from January 2022 to December 2026 and rendered at six fidelities, from a 3 MB unit-test fixture to a 5 GB full-scale corpus. 38,964,663 rows · 11 GB · 2,319 verification assertions, all passing. 100% synthetic. No real company… See the full description on the dataset page: https://huggingface.co/datasets/hemanthreddy901/nadora-global-industries.documenttabular-regression10K<n<100K0 likes722 downloads8d agoHugging Face04cutycat2000x /Global-LLMs-Replies Global LLMs Replies GPT-4o -> 74,644 rows mixtral-8x22b -> 13,129 rows claude-3-haiku -> 3,871 rows textquestion-answering10K<n<100K3 likes593 downloads2y agoHugging Face05apart-global-south-hack /remote_sensing_VQA_multilingual Remote Sensing VQA — Multilingual A multilingual counterfactual MCQ dataset built from remote sensing / satellite imagery. Each row contains a satellite image, two captions (original vs counterfactual), and a multiple-choice question probing whether a VLM follows the image or the misleading text. Languages Language Code Rows English en 50 Hindi hi 50 Urdu ur 50 Telugu te 50 Bahasa Indonesia id 50 Columns Column Type… See the full description on the dataset page: https://huggingface.co/datasets/apart-global-south-hack/remote_sensing_VQA_multilingual.imagequestion-answeringn<1K1 likes500 downloads3mo agoHugging Face06emperor-mew /global-censorship-index Voidly Global Censorship Index Real-time internet censorship measurements for 200 countries, based on 38,780,449+ OONI network probes. Dataset Description The Global Censorship Index provides country-level internet censorship scores derived from actual network measurements. Unlike annual expert assessments, this data updates daily. Key Statistics Countries covered: 200 Total measurements: 38,780,449 Severe censorship: 1 countries High censorship: 6… See the full description on the dataset page: https://huggingface.co/datasets/emperor-mew/global-censorship-index.tabulartext-classificationn<1K1 likes352 downloads2mo agoHugging Face07powertronglobal /powertron-global-permafrost-corpus Dataset Card: Powertron Global PermaFrost Corpus Important Disambiguation: This corpus documents PermaFrost® NMR, a trademarked HVAC efficiency treatment product. It contains HVAC/refrigeration efficiency data (chillers, RTUs, DX systems, refrigeration). This corpus has NO connection to geological permafrost (frozen ground), climate science, or Arctic research. The name "PermaFrost" is a product trademark reflecting thermal transfer properties, not a geological term.… See the full description on the dataset page: https://huggingface.co/datasets/powertronglobal/powertron-global-permafrost-corpus.tabulartext-generation1K<n<10K0 likes300 downloads2mo agoHugging Face08amayuelas /aya-global-exams-catalanCatalan exams for the Aya Global Exams. Original data and file available here: link Github Repo: link textquestion-answering1K<n<10K0 likes287 downloads2y agoHugging Face09apart-global-south-hack /counterfactual-pendulum-multilingual 📌 Dataset Summary When a Vision-Language Model (VLM) is given an image along with a text prompt containing contradictory or misleading information, how does it react? Does it rely on the visual evidence, succumb to textual bias, or honestly abstain when faced with unresolvable conflict? This dataset adapts the Counterfactual Pendulum scenario across two visual conflict dimensions: Angular (Angle): Conflict in the pendulum's angle of inclination. Light: Conflict in the light… See the full description on the dataset page: https://huggingface.co/datasets/apart-global-south-hack/counterfactual-pendulum-multilingual.imagequestion-answeringn<1K1 likes238 downloads3mo agoHugging Face10Kasher13 /prospire-synth-global-personas 🌍 Prospire Synth Global Personas The World's Largest Unified Synthetic Persona Database 512M+ records · 82 columns · 77+ countries · 39 languages · DuckDB-native 🎯 What Is This? Prospire Synth Global Personas is a unified, query-ready database of synthetic human personas built for AI agent simulations, market research, and cultural analysis. It merges 18 open-source datasets into a single coherent Parquet warehouse — partitioned, compressed, and… See the full description on the dataset page: https://huggingface.co/datasets/Kasher13/prospire-synth-global-personas.texttext-generation100M<n<1B1 likes206 downloads6mo agoHugging Face11metehan777 /global-seo-knowledgetexttext-generation1K<n<10K3 likes200 downloads1y agoHugging Face12tilikumotp /Global-Ocean-Science-Corpus 🌊 Global-Ocean-Science-Corpus (v2.0 Curated & Cleaned) A Highly Curated, Large-Scale Pre-Training & RAG Corpus for Deep Ocean Sciences, Marine Biology, and Oceanography Language Note: This dataset is a 100% English-language scientific corpus (language: "en") aggregating peer-reviewed literature, deep-sea exploration dossiers, and technical oceanographic reports from leading global marine institutes. Global-Ocean-Science-Corpus, derin okyanus bilimleri, deniz biyolojisi… See the full description on the dataset page: https://huggingface.co/datasets/tilikumotp/Global-Ocean-Science-Corpus.tabulartext-generation10K<n<100K1 likes103 downloads1d agoHugging Face13manuelcaccone /actuarial-global-glossary-multilingual 🤝 Connect with me on LinkedIn! Join the mission to make actuarial knowledge accessible worldwide Let's discuss how AI can transform professional education and break language barriers in finance! 🌍 Global Actuarial Glossary - Breaking Language Barriers in Finance 🚀 The World's Most Comprehensive Multilingual Actuarial Dataset Imagine: A brilliant actuarial student in Tokyo, a risk analyst in São Paulo, and an insurance executive… See the full description on the dataset page: https://huggingface.co/datasets/manuelcaccone/actuarial-global-glossary-multilingual.texttext-classification1K<n<10K0 likes98 downloads1y agoHugging Face14mariocedo /GlobalMedQA GlobalMedQA — A Standardized Multilingual Dataset for Assessing Medical Knowledge in LLMs Dataset Summary GlobalMedQA is a harmonized multilingual dataset of medical multiple-choice questions (MCQs) designed to benchmark large language models in the healthcare domain.It integrates exam questions from 14 countries and 13 languages, standardized into a unified schema with consistent metadata and specialty classification based on the European Union of Medical Specialists… See the full description on the dataset page: https://huggingface.co/datasets/mariocedo/GlobalMedQA.tabularquestion-answering100K<n<1M1 likes93 downloads11mo agoHugging Face15Mahadih534 /Global_Environment-Social-And-Governance-Data Global_Environment-Social-And-Governance Dataset This Dataset contains all verified and authorized Environment, Social and Governance Statistics data in the World Description I have collected all data from WORLD-Bank's Data Catalog and also shared this link in the data source section, this dataset is sutitable for various NLP tasks Data Source https://datacatalog.worldbank.org/ Dataset Card Authors Mahadi Hassan Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Global_Environment-Social-And-Governance-Data.tabularquestion-answering10K<n<100K1 likes56 downloads2y agoHugging Face16MaatAI /histoire-general-afrique-global-adaption This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. Svngoku/Histoire-General-Afrique-Global This dataset contains French-language text excerpts detailing the political, social, and economic history of Africa from the 16th to the 18th centuries. The content covers specific regions such as the Lower Guinea Coast and the Zambezi, discussing topics like ethnic migrations, kingdom formations, and trade dynamics. Each sample consists… See the full description on the dataset page: https://huggingface.co/datasets/MaatAI/histoire-general-afrique-global-adaption.textquestion-answering1K<n<10K1 likes56 downloads5mo agoHugging Face17cyankiwi /Global-MMLU-Lite-sample Global-MMLU-Lite — Sample Subset A small, fixed-size subset of CohereLabs/Global-MMLU-Lite intended for fast smoke-testing of multilingual MMLU evaluation pipelines. What this is 40 examples per language, 15 languages, 600 examples total. Strictly non-overlapping windows across languages: language i takes rows [i*40, i*40+40) from the per-language pool of test followed by dev (test=400, dev=200, pool=600). Concretely, languages 0–9 fall entirely within the original test… See the full description on the dataset page: https://huggingface.co/datasets/cyankiwi/Global-MMLU-Lite-sample.textquestion-answeringn<1K1 likes48 downloads5mo agoHugging Face18ai-for-good-lab /Global-MMLU-Lite Global MMLU-Lite — Human Translated Global MMLU-Lite is a multilingual evaluation benchmark for LLMs covering 18 languages. This dataset extends it with professional human translations for three additional low-resource languages that are not in the original: Chichewa (nya), Māori (mri), and Inuktitut (iku). Released as part of the BYOL: Bring Your Own Language Into LLMs project (paper). What's New The original Global MMLU-Lite by Cohere covers 18 languages: Arabic… See the full description on the dataset page: https://huggingface.co/datasets/ai-for-good-lab/Global-MMLU-Lite.textquestion-answering1K<n<10K1 likes47 downloads5mo agoHugging Face19sapiens-technology /global_mmlu_lite_pt 🌎 Global-MMLU Lite (Portuguese) A Focused Benchmark for Portuguese-Language Reasoning in Large Language Models Global-MMLU Lite (Portuguese) is a curated subset of the Global-MMLU Lite benchmark designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models in Portuguese, providing a diverse and computationally efficient collection of translated and adapted QA samples across domains such as general knowledge, science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_pt.textquestion-answeringn<1K0 likes41 downloads5mo agoHugging Face20amayuelas /aya-global-exams-spanishSpanish exams for the Aya Global Exams. Original data and file available here: link Github Repo: link textmultiple-choicen<1K0 likes36 downloads2y agoHugging Face21Nooxus-AI /NOO-Verified-Global-Entities 🌐 NOO-Verified-Global-Entities (NOO AI互联网商业实体验证归属) Official Data Infrastructure by Nooxus.com [ 🚀 ANGEL ROUND INVESTOR NOTICE / 天使轮国际融资公告 ] EN: Nooxus-AI is raising its Angel Round to scale nooxus.com — the world's first dedicated B2B Trading & Clearing Network for AI Agents. If your fund recognizes the trillion-dollar potential of building the "Visa / SWIFT network for the Agentic Web," powered by our production-ready 0.8ms RST signaling and Zero-Inbound stealth architecture… See the full description on the dataset page: https://huggingface.co/datasets/Nooxus-AI/NOO-Verified-Global-Entities.texttext-generation10K<n<100K1 likes36 downloads4mo agoHugging Face22sahilmaniyar888 /adaption-v5-global-employment-law-qa WorkRight V5 — Global Employment Law QA 1,751 employment-law reasoning examples across five jurisdictions, with a deterministic gold path and no language model anywhere in it. sha256: 2e828767c42a9f5ff53083509bacf6b1313e86027d418048376ac9441e25a456 — byte-identical to the artefact evaluated on Adaption. 1. Dataset Summary Every statutory answer here is produced by an executable rule that reads a fact scenario and returns a structured record — eligibility, amount… See the full description on the dataset page: https://huggingface.co/datasets/sahilmaniyar888/adaption-v5-global-employment-law-qa.textquestion-answering1K<n<10K1 likes30 downloads1mo agoHugging Face23Svngoku /histoire-general-afrique-global-adaption This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. Svngoku/Histoire-General-Afrique-Global This dataset contains French-language text excerpts detailing the political, social, and economic history of Africa from the 16th to the 18th centuries. The content covers specific regions such as the Lower Guinea Coast and the Zambezi, discussing topics like ethnic migrations, kingdom formations, and trade dynamics. Each sample consists… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/histoire-general-afrique-global-adaption.textquestion-answering1K<n<10K1 likes28 downloads5mo agoHugging Face24Srinivasmec26 /Educational-Flashcards-for-Global-Learners 1. Educational-Flashcards-for-Global-Learners/README.md Educational Flashcards Dataset Overview A comprehensive collection of 100 educational flashcards covering STEM, humanities, law, arts, and cultural topics. Curated with 70% Indian content, 25% European, and 5% other Asian perspectives to promote diverse knowledge representation. Dataset Structure { "input": "Text description", "output": { "type": "flashcards", "topic": "Subject name"… See the full description on the dataset page: https://huggingface.co/datasets/Srinivasmec26/Educational-Flashcards-for-Global-Learners.texttext-classificationn<1K1 likes26 downloads1y agoHugging Face25proxectonos /GlobalPIQA_gl GlobalPIQA_gl Related paper: Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures Dataset Summary GlobalPIQA_gl is the Galician subset of GlobalPIQA, a multilingual benchmark for evaluating physical commonsense reasoning across more than 100 languages and cultural contexts. It is intended as an evaluation resource for models that must choose the most plausible solution to a practical physical situation. The dataset follows the… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/GlobalPIQA_gl.textquestion-answeringn<1K0 likes26 downloads2mo agoHugging Face26gitrelief /law-global The Case-law, centralizing legal decisions for better use, a community Dataset. The Case-law Dataset is a comprehensive collection of legal decisons from various countries, centralized in a common format. This dataset aims to improve the development of legal AI models by providing a standardized, easily accessible corpus of global legal documents. Join us in our mission to make AI more accessible and understandable for the legal world, ensuring that the power of language… See the full description on the dataset page: https://huggingface.co/datasets/gitrelief/law-global.question-answering0 likes26 downloads2mo agoHugging Face27kelvinyelyen /tiny-aya-global-blindspots tiny-aya-global: Blind Spot Evaluation Model: CohereLabs/tiny-aya-global — ~3B parameter multilingual conversational model Dataset: kelvinyelyen/tiny-aya-global-blindspots Ten targeted probes designed to surface specific failure mechanisms, not aggregate accuracy. Each probe was run once under greedy decoding (temperature=0.0), then re-run 5x under sampling (temperature=0.7) to check whether each failure is a stable pattern or a one-off. Result: 7 clear failures, 1 pass, 1… See the full description on the dataset page: https://huggingface.co/datasets/kelvinyelyen/tiny-aya-global-blindspots.texttext-generationn<1K0 likes24 downloads2mo agoHugging Face28Mahadih534 /Global_Health-Nutrition-And-Population-Statistics Global_Health-Nutrition-And-Population-Statistics Dataset This Dataset contains all verified and authorized Health, Nutrition and Population Statistics data in the World Description I have collected all data from WORLD-Bank's Data Catalog and also shared this link in the data source section, this dataset is sutitable for various NLP tasks Data Source https://datacatalog.worldbank.org/ Dataset Card Authors Mahadi Hassan Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Global_Health-Nutrition-And-Population-Statistics.tabularquestion-answering100K<n<1M2 likes22 downloads2y agoHugging Face29amayuelas /aya-global-exams-basqueSpanish exams for the Aya Global Exams. Original data and file available here: link Github Repo: link textmultiple-choicen<1K0 likes19 downloads2y agoHugging Face30Owos /global-mmlu-lite Global MMLU Lite – Galician & Urdu Machine-translated Galician and Urdu subsets of the Global MMLU Lite benchmark. Dataset Description Global MMLU Lite is a culturally-aware, multilingual evaluation benchmark for large language models, covering multiple-choice questions across many academic subjects. This repository contains Galician (gl) and Urdu (ur) translations. This dataset was translated using Google Machine Translate. Splits Config… See the full description on the dataset page: https://huggingface.co/datasets/Owos/global-mmlu-lite.textquestion-answeringn<1K0 likes18 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.