datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.pii-masking-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.afrimmlu
Dataset Card for afrimmlu
Dataset Summary
AFRIMMLU is an evaluation dataset comprising translations of a subset of the MMLU dataset into 15 African languages.
It includes test sets across all 17 languages, maintaining an English and French subsets from the original MMLU dataset.
Languages
There are 17 languages available :
Dataset Structure
Data Instances
The examples look like this for English:
from datasets import load_dataset
data =… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrimmlu.pii-masking-400k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.open-pii-masking-500k-ai4privacy
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.master-dataset-all-V2
Master Dataset All V2
Google NQ Sequentially Sharded Dataset.
100k-corpus-2026
MAST 100K Corpus 2026
This dataset contains the fixed English document corpus used for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages by retrieving English evidence and producing short, correct English answers.
This corpus is copied from BrowseComp-Plus, a benchmark for Deep-Research systems that isolates the effect of the retriever and the LLM agent to… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/100k-corpus-2026.MasterMind
Dataset Card for MasterMind
English | 简体中文(Simplified Chinese)
Dataset Description
Dataset Summary
This dataset contains the expert dataset for the Doudizhu and Go tasks proposed in MasterMind. In summary, this dataset uses a QA format, with the question part providing the current state of the game; the answer part provides the corresponding game-playing strategy and the logic behind adopting this strategy. The dataset encodes all the above information in… See the full description on the dataset page: https://huggingface.co/datasets/OpenDILabCommunity/MasterMind.uhura-truthfulqa
Dataset Card for Uhura-TruthfulQA
Dataset Summary
TruthfulQA is a widely recognized safety benchmark designed to measure the truthfulness of language model outputs across 38 categories, including health, law, finance, and politics. The English version of the benchmark originates from TruthfulQA: Measuring How Models Mimic Human Falsehoods (Lin et al., 2022) and consists of 817 questions in both multiple-choice and generation formats, targeting common misconceptions and… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/uhura-truthfulqa.uhura-arc-easy
Dataset Card for Uhura-Arc-Easy
Dataset Summary
Uhura-ARC-Easy is a widely recognized scientific question answering benchmark composed of multiple-choice science questions derived from grade-school examinations that test various styles of knowledge and reasoning.
The original English version of the benchmark originates from Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge (Clark et al., 2018) and is divided into "Challenge" and "Easy"… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/uhura-arc-easy.pii-masking-200k
Purpose and Features
World's largest open source privacy dataset.
The purpose of the dataset is to train models to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs.
The example texts have 54 PII classes (types of sensitive data), targeting 229 discussion subjects / use cases split across business, education, psychology and legal fields, and 5 interactions styles (e.g. casual conversation, formal document, emails… See the full description on the dataset page: https://huggingface.co/datasets/Isotonic/pii-masking-200k.indic-queries-2026
MAST Indic Queries 2026
This dataset contains the Indic query set for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages.
MAST builds on BrowseComp-Plus (ACL 2026), a reproducible and verifiable extension of BrowseComp with challenging English queries, a verified English corpus of roughly 100K web-sourced documents, and human judgments. In the 2026 MAST… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/indic-queries-2026.multilingual-queries-2026
MAST Multilingual Queries 2026
This dataset contains the multilingual query set for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages.
MAST builds on BrowseComp-Plus (ACL 2026), a reproducible and verifiable extension of BrowseComp with challenging English queries, a verified English corpus of roughly 100K web-sourced documents, and human judgments. In the… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/multilingual-queries-2026.long-horizon
ToolGym Long-Horizon Dataset
Dataset Description
This dataset contains long-horizon trajectories and evaluations for the ToolGym benchmark.
Dataset Structure
long-horizon/
├── traj/ # Agent trajectories (JSONL format)
│ ├── gpt-5.2/
│ │ ├── pass@1.jsonl
│ │ ├── pass@2.jsonl
│ │ └── pass@3.jsonl
│ ├── claude-opus-4.5/
│ └── ...
└── eval/ # Evaluation results (JSONL format)
├── claude-opus-4.5/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/masculine/long-horizon.AtomWorldBench
AtomWorldBench
AtomWorldBench is a benchmark and dataset for evaluating the ability of Large Language Models (LLMs) and agents to perform 3D crystal structure manipulation from natural language instructions.
Given an input crystal structure in CIF format and a textual instruction, the model must generate the resulting crystal structure after applying the requested modification.
The dataset is released alongside the AtomWorld benchmark framework and is intended for:
Benchmarking… See the full description on the dataset page: https://huggingface.co/datasets/Master-AI-Lab/AtomWorldBench.cybersec-master-dataset
Cybersecurity Master Instruction Dataset
Overview
A large-scale cybersecurity instruction-tuning dataset in ShareGPT conversational format,
assembled from multiple authoritative open sources and deduplicated.
At 1,807,941 deduplicated records, this appears to be one of the larger cybersecurity
LLM fine-tuning / instruction-style datasets on Hugging Face, and likely among the larger
broad vulnerability-intelligence corpora in conversational/instruction format. It is… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/cybersec-master-dataset.OpenMath-GSM8K-masked
OpenMath GSM8K Masked
We release a masked version of the GSM8K solutions.
This data can be used to aid synthetic generation of additional solutions for GSM8K dataset
as it is much less likely to lead to inconsistent reasoning compared to using
the original solutions directly.
This dataset was used to construct OpenMathInstruct-1:
a math instruction tuning dataset with 1.8M problem-solution pairs
generated using permissively licensed Mixtral-8x7B model.
For details of how the masked… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMath-GSM8K-masked.k-beauty-ai-citation-dataset
K-Beauty AI Citation Dataset
Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research.
Canonical source: https://kbeautyanswers.com/dataset/
License: CC BY 4.0
Maintainer: K-Beauty Answers (site)
Initial release: 2026-05-23
What's in it
128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.mASNQ
Dataset Description
mASNQ is a translated version of ASNQ which is an AS2 dataset created by adapting the Natural Question corpus from Machine Reading (MR) to the AS2 task.
The dataset has been translated into five European languages: French, German, Italian, Portuguese, and Spanish, as described in this paper: Datasets for Multilingual Answer Sentence Selection.
Splits:
For each language (English, French, German, Italian, Portuguese, and Spanish), we provide:… See the full description on the dataset page: https://huggingface.co/datasets/matteogabburo/mASNQ.sai-mash
Multilingual Audits: Structured & Harmonized
MASH is a dataset of Supreme Audit Institution (SAI) reports harmonized to a common language and format. SAIs publish their work in national languages and with varying structures, making cross-country analysis difficult. MASH resolves this by processing each report through a standardized pipeline that produces English summaries, structured metadata, and controlled-vocabulary tags — enabling researchers, auditors, and developers to… See the full description on the dataset page: https://huggingface.co/datasets/Riksrevisjonen/sai-mash.afriqaAfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages
AfriQA is the first cross-lingual question answering (QA) dataset with a focus on African languages.
The dataset includes over 12,000 XOR QA examples across 10 African languages, making it an invaluable resource for developing more equitable QA technology.master-dataset-all-V1
Master Dataset All V2 (Part-1 Shards)
This repository contains sequentially sharded parts extracted from Google's Natural Questions dataset to optimize training and ingestion loops for LLM fine-tuning.
Dataset Structure
Format: JSON Lines (.jsonl)
Shards Uploaded: train-00000.jsonl to train-00325.jsonl (Part-1)
Data Configuration: Out-of-the-box support for datasets loader.
Generated and uploaded sequentially via RunPod pipeline.
project-madurai-booksProject Madurai Books Text Dataset
This dataset card aims to convert the Tamil books available on the Project Madurai website to the HF dataset. It has been scrapped from Project Madurai Website.
Dataset Details
You can see a table above called "Meta Data", which is just an info table.
You can't able to preview the "Source Data" table, due to it being about 300MB.
[Don't open the Dataset in Excel It will lead to a crash of the OS instead open it using Python in pandas or… See the full description on the dataset page: https://huggingface.co/datasets/mastergokul/project-madurai-books.OpenMath-MATH-masked
OpenMath GSM8K Masked
We release a masked version of the MATH solutions.
This data can be used to aid synthetic generation of additional solutions for MATH dataset
as it is much less likely to lead to inconsistent reasoning compared to using
the original solutions directly.
This dataset was used to construct OpenMathInstruct-1:
a math instruction tuning dataset with 1.8M problem-solution pairs
generated using permissively licensed Mixtral-8x7B model.
For details of how the masked… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMath-MATH-masked.MasterMind
Dataset Card for MasterMind
English | 简体中文(Simplified Chinese)
Dataset Description
Dataset Summary
This dataset contains the expert dataset for the Doudizhu and Go tasks proposed in MasterMind. In summary, this dataset uses a QA format, with the question part providing the current state of the game; the answer part provides the corresponding game-playing strategy and the logic behind adopting this strategy. The dataset encodes all the above information in… See the full description on the dataset page: https://huggingface.co/datasets/JoanhLan/MasterMind.pii-masking-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ASR2005Bluesnow/pii-masking-200k.short-horizon
ToolGym Short-Horizon Dataset
Dataset Description
This dataset contains short-horizon trajectories and evaluations for the ToolGym benchmark.
Dataset Structure
short-horizon/
├── traj/ # Agent trajectories (JSONL format)
│ ├── claude-3.5/
│ │ ├── pass@1.jsonl
│ │ ├── pass@2.jsonl
│ │ └── pass@3.jsonl
│ ├── deepseek-v3.2/
│ └── ...
└── eval/ # Evaluation results (JSONL format)
├── claude-3.5/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/masculine/short-horizon.SwS-Demo-Dataset
Dataset Card for SwS-Demo-Dataset
[🌐 Website] •
[🤗 Demo Dataset] •
[📜 Paper] •
[🐱 GitHub] •
[🐦 Twitter] •
[📕 Rednote]
This dataset is a demo set of synthetic problems generated by SwS, comprising 500 samples for each model and category. The full dataset and model are currently under review by Microsoft and will be released once approved.
Data Loading
from datasets import load_dataset
dataset = load_dataset("MasterVito/SwS-Demo-Dataset")
Data… See the full description on the dataset page: https://huggingface.co/datasets/MasterVito/SwS-Demo-Dataset.AI_Mastery_Foundation_Curriculum
FOUNDATION DATASET
AI Mastery Foundation
Curriculum
A premium foundation layer for knowledge, reasoning, preference,
reward, benchmark, and agentic tool-use training.
Hugging Face-ready Parquet package
AI Mastery Foundation Curriculum
A premium staged foundation dataset for building models with a cleaner first layer of academic… See the full description on the dataset page: https://huggingface.co/datasets/ayjays132/AI_Mastery_Foundation_Curriculum.afriqa-gold-passagesAfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages
AfriQA is the first cross-lingual question-answering (QA) dataset with a focus on African languages.
The dataset includes over 12,000 XOR QA examples across 10 African languages, making it an invaluable resource for developing more equitable QA technology.
