CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai4privacy /pii-masking-300k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.texttext-classification100K<n<1M116 likes4.3k downloads4mo agoHugging Face02ai4privacy /pii-masking-200k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.texttext-classification100K<n<1M127 likes3.5k downloads4mo agoHugging Face03masakhane /afrimmlu Dataset Card for afrimmlu Dataset Summary AFRIMMLU is an evaluation dataset comprising translations of a subset of the MMLU dataset into 15 African languages. It includes test sets across all 17 languages, maintaining an English and French subsets from the original MMLU dataset. Languages There are 17 languages available : Dataset Structure Data Instances The examples look like this for English: from datasets import load_dataset data =… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/afrimmlu.textquestion-answering10K<n<100K12 likes1.7k downloads1y agoHugging Face04ai4privacy /pii-masking-400k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.texttext-classification100K<n<1M63 likes1.3k downloads4mo agoHugging Face05ai4privacy /open-pii-masking-500k-ai4privacy 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.texttext-classification100K<n<1M27 likes1.2k downloads4mo agoHugging Face06anupbth1 /master-dataset-all-V2 Master Dataset All V2 Google NQ Sequentially Sharded Dataset. question-answering0 likes752 downloads4mo agoHugging Face07mast-benchmark /100k-corpus-2026 MAST 100K Corpus 2026 This dataset contains the fixed English document corpus used for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages by retrieving English evidence and producing short, correct English answers. This corpus is copied from BrowseComp-Plus, a benchmark for Deep-Research systems that isolates the effect of the retriever and the LLM agent to… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/100k-corpus-2026.textquestion-answering100K<n<1M0 likes482 downloads2mo agoHugging Face08OpenDILabCommunity /MasterMind Dataset Card for MasterMind English | 简体中文(Simplified Chinese) Dataset Description Dataset Summary This dataset contains the expert dataset for the Doudizhu and Go tasks proposed in MasterMind. In summary, this dataset uses a QA format, with the question part providing the current state of the game; the answer part provides the corresponding game-playing strategy and the logic behind adopting this strategy. The dataset encodes all the above information in… See the full description on the dataset page: https://huggingface.co/datasets/OpenDILabCommunity/MasterMind.textquestion-answering100K<n<1M6 likes358 downloads2y agoHugging Face09masakhane /uhura-truthfulqa Dataset Card for Uhura-TruthfulQA Dataset Summary TruthfulQA is a widely recognized safety benchmark designed to measure the truthfulness of language model outputs across 38 categories, including health, law, finance, and politics. The English version of the benchmark originates from TruthfulQA: Measuring How Models Mimic Human Falsehoods (Lin et al., 2022) and consists of 817 questions in both multiple-choice and generation formats, targeting common misconceptions and… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/uhura-truthfulqa.textmultiple-choice10K<n<100K2 likes320 downloads2y agoHugging Face10masakhane /uhura-arc-easy Dataset Card for Uhura-Arc-Easy Dataset Summary Uhura-ARC-Easy is a widely recognized scientific question answering benchmark composed of multiple-choice science questions derived from grade-school examinations that test various styles of knowledge and reasoning. The original English version of the benchmark originates from Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge (Clark et al., 2018) and is divided into "Challenge" and "Easy"… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/uhura-arc-easy.textmultiple-choice1K<n<10K1 likes319 downloads2y agoHugging Face11Isotonic /pii-masking-200k Purpose and Features World's largest open source privacy dataset. The purpose of the dataset is to train models to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs. The example texts have 54 PII classes (types of sensitive data), targeting 229 discussion subjects / use cases split across business, education, psychology and legal fields, and 5 interactions styles (e.g. casual conversation, formal document, emails… See the full description on the dataset page: https://huggingface.co/datasets/Isotonic/pii-masking-200k.texttext-classification100K<n<1M9 likes300 downloads3y agoHugging Face12mast-benchmark /indic-queries-2026 MAST Indic Queries 2026 This dataset contains the Indic query set for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages. MAST builds on BrowseComp-Plus (ACL 2026), a reproducible and verifiable extension of BrowseComp with challenging English queries, a verified English corpus of roughly 100K web-sourced documents, and human judgments. In the 2026 MAST… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/indic-queries-2026.textquestion-answeringn<1K1 likes275 downloads2mo agoHugging Face13mast-benchmark /multilingual-queries-2026 MAST Multilingual Queries 2026 This dataset contains the multilingual query set for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages. MAST builds on BrowseComp-Plus (ACL 2026), a reproducible and verifiable extension of BrowseComp with challenging English queries, a verified English corpus of roughly 100K web-sourced documents, and human judgments. In the… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/multilingual-queries-2026.textquestion-answeringn<1K1 likes268 downloads1mo agoHugging Face14masculine /long-horizon ToolGym Long-Horizon Dataset Dataset Description This dataset contains long-horizon trajectories and evaluations for the ToolGym benchmark. Dataset Structure long-horizon/ ├── traj/ # Agent trajectories (JSONL format) │ ├── gpt-5.2/ │ │ ├── pass@1.jsonl │ │ ├── pass@2.jsonl │ │ └── pass@3.jsonl │ ├── claude-opus-4.5/ │ └── ... └── eval/ # Evaluation results (JSONL format) ├── claude-opus-4.5/ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/masculine/long-horizon.text-generation1K<n<10K0 likes232 downloads9mo agoHugging Face15Master-AI-Lab /AtomWorldBench AtomWorldBench AtomWorldBench is a benchmark and dataset for evaluating the ability of Large Language Models (LLMs) and agents to perform 3D crystal structure manipulation from natural language instructions. Given an input crystal structure in CIF format and a textual instruction, the model must generate the resulting crystal structure after applying the requested modification. The dataset is released alongside the AtomWorld benchmark framework and is intended for: Benchmarking… See the full description on the dataset page: https://huggingface.co/datasets/Master-AI-Lab/AtomWorldBench.textquestion-answering10K<n<100K1 likes206 downloads4mo agoHugging Face16Voidreaper2026 /cybersec-master-dataset Cybersecurity Master Instruction Dataset Overview A large-scale cybersecurity instruction-tuning dataset in ShareGPT conversational format, assembled from multiple authoritative open sources and deduplicated. At 1,807,941 deduplicated records, this appears to be one of the larger cybersecurity LLM fine-tuning / instruction-style datasets on Hugging Face, and likely among the larger broad vulnerability-intelligence corpora in conversational/instruction format. It is… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/cybersec-master-dataset.texttext-generation1M<n<10M4 likes195 downloads5mo agoHugging Face17nvidia /OpenMath-GSM8K-masked OpenMath GSM8K Masked We release a masked version of the GSM8K solutions. This data can be used to aid synthetic generation of additional solutions for GSM8K dataset as it is much less likely to lead to inconsistent reasoning compared to using the original solutions directly. This dataset was used to construct OpenMathInstruct-1: a math instruction tuning dataset with 1.8M problem-solution pairs generated using permissively licensed Mixtral-8x7B model. For details of how the masked… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMath-GSM8K-masked.textquestion-answering1K<n<10K12 likes191 downloads3y agoHugging Face18k-master /k-beauty-ai-citation-dataset K-Beauty AI Citation Dataset Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research. Canonical source: https://kbeautyanswers.com/dataset/ License: CC BY 4.0 Maintainer: K-Beauty Answers (site) Initial release: 2026-05-23 What's in it 128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.texttext-classificationn<1K0 likes181 downloads3mo agoHugging Face19matteogabburo /mASNQ Dataset Description mASNQ is a translated version of ASNQ which is an AS2 dataset created by adapting the Natural Question corpus from Machine Reading (MR) to the AS2 task. The dataset has been translated into five European languages: French, German, Italian, Portuguese, and Spanish, as described in this paper: Datasets for Multilingual Answer Sentence Selection. Splits: For each language (English, French, German, Italian, Portuguese, and Spanish), we provide:… See the full description on the dataset page: https://huggingface.co/datasets/matteogabburo/mASNQ.tabularquestion-answering10M<n<100M0 likes156 downloads2y agoHugging Face20Riksrevisjonen /sai-mash Multilingual Audits: Structured & Harmonized MASH is a dataset of Supreme Audit Institution (SAI) reports harmonized to a common language and format. SAIs publish their work in national languages and with varying structures, making cross-country analysis difficult. MASH resolves this by processing each report through a standardized pipeline that produces English summaries, structured metadata, and controlled-vocabulary tags — enabling researchers, auditors, and developers to… See the full description on the dataset page: https://huggingface.co/datasets/Riksrevisjonen/sai-mash.tabularquestion-answering1K<n<10K2 likes151 downloads15d agoHugging Face21masakhane /afriqaAfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages AfriQA is the first cross-lingual question answering (QA) dataset with a focus on African languages. The dataset includes over 12,000 XOR QA examples across 10 African languages, making it an invaluable resource for developing more equitable QA technology.textquestion-answering10K<n<100K11 likes150 downloads3y agoHugging Face22anupbth1 /master-dataset-all-V1 Master Dataset All V2 (Part-1 Shards) This repository contains sequentially sharded parts extracted from Google's Natural Questions dataset to optimize training and ingestion loops for LLM fine-tuning. Dataset Structure Format: JSON Lines (.jsonl) Shards Uploaded: train-00000.jsonl to train-00325.jsonl (Part-1) Data Configuration: Out-of-the-box support for datasets loader. Generated and uploaded sequentially via RunPod pipeline. question-answering0 likes129 downloads4mo agoHugging Face23mastergokul /project-madurai-booksProject Madurai Books Text Dataset This dataset card aims to convert the Tamil books available on the Project Madurai website to the HF dataset. It has been scrapped from Project Madurai Website. Dataset Details You can see a table above called "Meta Data", which is just an info table. You can't able to preview the "Source Data" table, due to it being about 300MB. [Don't open the Dataset in Excel It will lead to a crash of the OS instead open it using Python in pandas or… See the full description on the dataset page: https://huggingface.co/datasets/mastergokul/project-madurai-books.tabulartext-classification1K<n<10K0 likes126 downloads2y agoHugging Face24nvidia /OpenMath-MATH-masked OpenMath GSM8K Masked We release a masked version of the MATH solutions. This data can be used to aid synthetic generation of additional solutions for MATH dataset as it is much less likely to lead to inconsistent reasoning compared to using the original solutions directly. This dataset was used to construct OpenMathInstruct-1: a math instruction tuning dataset with 1.8M problem-solution pairs generated using permissively licensed Mixtral-8x7B model. For details of how the masked… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMath-MATH-masked.textquestion-answering1K<n<10K9 likes121 downloads3y agoHugging Face25JoanhLan /MasterMind Dataset Card for MasterMind English | 简体中文(Simplified Chinese) Dataset Description Dataset Summary This dataset contains the expert dataset for the Doudizhu and Go tasks proposed in MasterMind. In summary, this dataset uses a QA format, with the question part providing the current state of the game; the answer part provides the corresponding game-playing strategy and the logic behind adopting this strategy. The dataset encodes all the above information in… See the full description on the dataset page: https://huggingface.co/datasets/JoanhLan/MasterMind.textquestion-answering100K<n<1M0 likes78 downloads9mo agoHugging Face26ASR2005Bluesnow /pii-masking-200k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ASR2005Bluesnow/pii-masking-200k.texttext-classification100K<n<1M0 likes78 downloads24d agoHugging Face27masculine /short-horizon ToolGym Short-Horizon Dataset Dataset Description This dataset contains short-horizon trajectories and evaluations for the ToolGym benchmark. Dataset Structure short-horizon/ ├── traj/ # Agent trajectories (JSONL format) │ ├── claude-3.5/ │ │ ├── pass@1.jsonl │ │ ├── pass@2.jsonl │ │ └── pass@3.jsonl │ ├── deepseek-v3.2/ │ └── ... └── eval/ # Evaluation results (JSONL format) ├── claude-3.5/ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/masculine/short-horizon.text-generation1K<n<10K0 likes77 downloads9mo agoHugging Face28MasterVito /SwS-Demo-Dataset Dataset Card for SwS-Demo-Dataset [🌐 Website] • [🤗 Demo Dataset] • [📜 Paper] • [🐱 GitHub] • [🐦 Twitter] • [📕 Rednote] This dataset is a demo set of synthetic problems generated by SwS, comprising 500 samples for each model and category. The full dataset and model are currently under review by Microsoft and will be released once approved. Data Loading from datasets import load_dataset dataset = load_dataset("MasterVito/SwS-Demo-Dataset") Data… See the full description on the dataset page: https://huggingface.co/datasets/MasterVito/SwS-Demo-Dataset.textquestion-answering10K<n<100K2 likes74 downloads1y agoHugging Face29ayjays132 /AI_Mastery_Foundation_Curriculum FOUNDATION DATASET AI Mastery Foundation Curriculum A premium foundation layer for knowledge, reasoning, preference, reward, benchmark, and agentic tool-use training. Hugging Face-ready Parquet package AI Mastery Foundation Curriculum A premium staged foundation dataset for building models with a cleaner first layer of academic… See the full description on the dataset page: https://huggingface.co/datasets/ayjays132/AI_Mastery_Foundation_Curriculum.texttext-generation10K<n<100K1 likes70 downloads4mo agoHugging Face30masakhane /afriqa-gold-passagesAfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages AfriQA is the first cross-lingual question-answering (QA) dataset with a focus on African languages. The dataset includes over 12,000 XOR QA examples across 10 African languages, making it an invaluable resource for developing more equitable QA technology.question-answering10K<n<100K6 likes66 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.