CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai4privacy /pii-masking-300k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.texttext-classification100K<n<1M116 likes4.3k downloads4mo agoHugging Face02ai4privacy /pii-masking-200k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.texttext-classification100K<n<1M127 likes3.5k downloads4mo agoHugging Face03ai4privacy /pii-masking-openpii-1.5m OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension) 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Overview The OpenPII 1.5M dataset extends OpenPII 1M with a new Asia Pacific corpus, bringing global coverage to 30 languages across Europe, Americas, and Asia Pacific. This is the flagship release of the PII-Masking-3M family, the world's largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.texttoken-classification1M<n<10M21 likes2.9k downloads4mo agoHugging Face04DATA-MASK /FineWeb-Mask FineWeb-Mask 📜 DATAMASK Paper | 💻 GitHub Repository | 📦 Fineweb-Mask Dataset 📚 Introduction FineWeb-Mask is a 1.5 trillion token, high-efficiency pre-training dataset curated using the DATAMASK framework. Developed by the ByteDance Seed team, DATAMASK addresses the fundamental tension in large-scale data selection: the trade-off between high quality and high diversity. By modeling data selection as a Mask Learning problem, we provide a derivative of the original… See the full description on the dataset page: https://huggingface.co/datasets/DATA-MASK/FineWeb-Mask.text-generationn>1T6 likes2.4k downloads8mo agoHugging Face05gretelai /gretel-pii-masking-en-v1 Gretel Synthetic Domain-Specific Documents Dataset (English) This dataset is a synthetically generated collection of documents enriched with Personally Identifiable Information (PII) and Protected Health Information (PHI) entities spanning multiple domains. Created using Gretel Navigator with mistral-nemo-2407 as the backend model, it is specifically designed for fine-tuning Gliner models. The dataset contains document passages featuring PII/PHI entities from a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-pii-masking-en-v1.texttext-classification10K<n<100K46 likes1.7k downloads9mo agoHugging Face06ai4privacy /pii-masking-openpii-1m OpenPII 1M — Multilingual PII Masking Dataset Overview The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types. Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.texttoken-classification1M<n<10M15 likes1.5k downloads6mo agoHugging Face07ai4privacy /pii-masking-400k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.texttext-classification100K<n<1M63 likes1.3k downloads4mo agoHugging Face08ai4privacy /open-pii-masking-500k-ai4privacy 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.texttext-classification100K<n<1M27 likes1.2k downloads4mo agoHugging Face09Parsannazari12 /cybersecurity-master-dataset Cybersecurity Master Dataset Unified and deduplicated cybersecurity SFT dataset containing CTF solutions, CVE analyses, vulnerability patches, and Python coding instructions. texttext-generation100K<n<1M3 likes1.1k downloads24d agoHugging Face10Brainquiver /general-master-en-202608 General · Master · English · 2026-08 English pretraining text, assembled from three public sources, cleaned with one character-level cleaner, and filtered for repetition. 109,337,531 documents and 468,064,046,462 characters. Composition Config Documents Characters What it is fineweb-edu-dedup 65,010,430 297,544,916,118 Web text an educational classifier kept cosmopedia-v2 38,591,146 144,011,993,012 Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.tabulartext-generation100M<n<1B1 likes983 downloads27d agoHugging Face11MasahiroKaneko /eagle Eagle 🦅: Ethical Dataset Given from Real Interactions Introduction This repository contains the Eagle dataset, which is an ethical dataset of real interactions between humans and ChatGPT. This dataset is created to evaluate social bias, opinion bias, toxic language, and morality in Large Language Models (LLMs). If you use the Eagle dataset in your research, please cite the following: @inproceedings{Eagle:arxiv:2024, title={Eagle: Ethical Dataset Given from Real… See the full description on the dataset page: https://huggingface.co/datasets/MasahiroKaneko/eagle.tabulartext-generation100K<n<1M4 likes390 downloads3y agoHugging Face12masakhane /uhura-truthfulqa Dataset Card for Uhura-TruthfulQA Dataset Summary TruthfulQA is a widely recognized safety benchmark designed to measure the truthfulness of language model outputs across 38 categories, including health, law, finance, and politics. The English version of the benchmark originates from TruthfulQA: Measuring How Models Mimic Human Falsehoods (Lin et al., 2022) and consists of 817 questions in both multiple-choice and generation formats, targeting common misconceptions and… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/uhura-truthfulqa.textmultiple-choice10K<n<100K2 likes320 downloads2y agoHugging Face13Isotonic /pii-masking-200k Purpose and Features World's largest open source privacy dataset. The purpose of the dataset is to train models to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs. The example texts have 54 PII classes (types of sensitive data), targeting 229 discussion subjects / use cases split across business, education, psychology and legal fields, and 5 interactions styles (e.g. casual conversation, formal document, emails… See the full description on the dataset page: https://huggingface.co/datasets/Isotonic/pii-masking-200k.texttext-classification100K<n<1M9 likes300 downloads3y agoHugging Face14nielsr /MS-GPT-MassSpecGym MS-GPT MassSpecGym benchmark data This repository contains the MassSpecGym benchmark files and labels released with MS-GPT: Rethinking MS/MS De Novo Structure Elucidation as Spectrum-Induced Posterior Querying of a Molecule-Language Model. Provenance and attribution These files are mirrored from the authors' released MS-GPT asset bundle with their explicit authorization. Original repository: https://github.com/VIKI623/MS-GPT Author's Hugging Face profile:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/MS-GPT-MassSpecGym.texttext-generation10K<n<100K0 likes296 downloads2mo agoHugging Face15masculine /long-horizon ToolGym Long-Horizon Dataset Dataset Description This dataset contains long-horizon trajectories and evaluations for the ToolGym benchmark. Dataset Structure long-horizon/ ├── traj/ # Agent trajectories (JSONL format) │ ├── gpt-5.2/ │ │ ├── pass@1.jsonl │ │ ├── pass@2.jsonl │ │ └── pass@3.jsonl │ ├── claude-opus-4.5/ │ └── ... └── eval/ # Evaluation results (JSONL format) ├── claude-opus-4.5/ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/masculine/long-horizon.text-generation1K<n<10K0 likes232 downloads9mo agoHugging Face16Salesforce /MASBench 🎼 MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks This is the proposed MAS evaluation data used in the recipe described in our paper:📄 MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks For more details, please check the following resources: 🌐 Project Page: https://mas-orchestra.salesforceresearch.ai/mas_r1/index.html 📚 Live… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/MASBench.texttext-generation10K<n<100K3 likes227 downloads4mo agoHugging Face17masakhane /AfriADRtexttext-generation10K<n<100K2 likes222 downloads2y agoHugging Face18PiTtawat9 /MS-GPT-MassSpecGym MS-GPT MassSpecGym benchmark data This repository contains the MassSpecGym benchmark files and labels released with MS-GPT: Rethinking MS/MS De Novo Structure Elucidation as Spectrum-Induced Posterior Querying of a Molecule-Language Model. Provenance and attribution These files are mirrored from the authors' released MS-GPT asset bundle with their explicit authorization. Original repository: https://github.com/VIKI623/MS-GPT Author's Hugging Face profile:… See the full description on the dataset page: https://huggingface.co/datasets/PiTtawat9/MS-GPT-MassSpecGym.texttext-generation10K<n<100K0 likes214 downloads2mo agoHugging Face19masamasa4 /drawio-xmltext-generation0 likes205 downloads1y agoHugging Face20Voidreaper2026 /cybersec-master-dataset Cybersecurity Master Instruction Dataset Overview A large-scale cybersecurity instruction-tuning dataset in ShareGPT conversational format, assembled from multiple authoritative open sources and deduplicated. At 1,807,941 deduplicated records, this appears to be one of the larger cybersecurity LLM fine-tuning / instruction-style datasets on Hugging Face, and likely among the larger broad vulnerability-intelligence corpora in conversational/instruction format. It is… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/cybersec-master-dataset.texttext-generation1M<n<10M4 likes195 downloads5mo agoHugging Face21nvidia /OpenMath-GSM8K-masked OpenMath GSM8K Masked We release a masked version of the GSM8K solutions. This data can be used to aid synthetic generation of additional solutions for GSM8K dataset as it is much less likely to lead to inconsistent reasoning compared to using the original solutions directly. This dataset was used to construct OpenMathInstruct-1: a math instruction tuning dataset with 1.8M problem-solution pairs generated using permissively licensed Mixtral-8x7B model. For details of how the masked… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMath-GSM8K-masked.textquestion-answering1K<n<10K12 likes191 downloads3y agoHugging Face22massines3a /assistant-axis-vectors Assistant Axis Vectors for gemma-3-27b-it This dataset contains pre-computed role vectors and the assistant axis for gemma-3-27b-it. Overview These vectors were computed using the methodology from the paper "The Assistant Axis" by Christina Lu et al. The vectors can be used for activation steering to control model behavior along the "assistant-like" to "role-playing" spectrum. Contents gemma-3-27b-it/assistant_axis.pt - The computed assistant axis (principal… See the full description on the dataset page: https://huggingface.co/datasets/massines3a/assistant-axis-vectors.texttext-generationn<1K0 likes177 downloads8mo agoHugging Face23thepowerfuldeez /massive-yt-edu-queue Massive YouTube Educational Video Queue Full metadata and content classification for 4,489,228 YouTube educational videos totaling 3,975,157 hours. Description This dataset contains metadata, content categorization, and license risk assessment for ~4.5M YouTube videos identified as potentially educational. It serves as the discovery and processing queue for the massive-yt-edu-transcriptions project, which aims to create the world's largest open educational transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-queue.tabularautomatic-speech-recognition1M<n<10M1 likes146 downloads7mo agoHugging Face24ai4privacy /openpii-masking-micro-100k OpenPII Micro: Multilingual PII Masking Sample A micro-sized stratified sample of OpenPII 1.5M, perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every label that exists in the parent dataset is represented in proportion. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format License 100,000 90,000 10,000 19… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-micro-100k.texttoken-classification100K<n<1M0 likes133 downloads4mo agoHugging Face25nvidia /OpenMath-MATH-masked OpenMath GSM8K Masked We release a masked version of the MATH solutions. This data can be used to aid synthetic generation of additional solutions for MATH dataset as it is much less likely to lead to inconsistent reasoning compared to using the original solutions directly. This dataset was used to construct OpenMathInstruct-1: a math instruction tuning dataset with 1.8M problem-solution pairs generated using permissively licensed Mixtral-8x7B model. For details of how the masked… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMath-MATH-masked.textquestion-answering1K<n<10K9 likes121 downloads3y agoHugging Face26masharma /convolearn ConvoLearn A dataset of tutor-student conversations demonstrating dialogic (knowledge-building) pedagogies. What's in here 2,134 dialogues between teachers and a simulated 7th-grade student discussing middle school Earth Science. Each conversation demonstrates one of six knowledge-building dimensions: cognitive engagement, formative assessment, accountability, cultural responsiveness, metacognition, or power dynamics. The teachers were real educators (323 credentialed… See the full description on the dataset page: https://huggingface.co/datasets/masharma/convolearn.tabulartext-generation1K<n<10K2 likes112 downloads6mo agoHugging Face27ai4privacy /pii-masking-mini-10k PII Masking Mini: Multilingual Sample A mini-sized stratified sample of pii-masking-openpii-1.5m, the flagship release of the PII-Masking-3M family. Sampled proportionally by (source_dataset, language) so every locale and label gets representation. Asia Pacific rows appear first. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-mini-10k.texttoken-classification1K<n<10K0 likes110 downloads4mo agoHugging Face28ai4privacy /pii-masking-micro-100k PII Masking Micro: Multilingual Sample A micro-sized stratified sample of pii-masking-openpii-1.5m, the flagship release of the PII-Masking-3M family. Sampled proportionally by (source_dataset, language) so every locale and label gets representation. Asia Pacific rows appear first. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-micro-100k.texttoken-classification10K<n<100K0 likes109 downloads4mo agoHugging Face29ai4privacy /openpii-masking-nano-1k OpenPII Nano: Multilingual PII Masking Sample A nano-sized stratified sample of OpenPII 1.5M, perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every label that exists in the parent dataset is represented in proportion. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format License 1,000 900 100 19 30 37 7… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-nano-1k.texttoken-classification1K<n<10K3 likes107 downloads4mo agoHugging Face30Feng613 /MASS-EX MASS-EX: Expert-Annotated Dataset for Interpretable Sleep Staging 中文版 Associated Paper:Guifeng Deng, Pan Wang, Jiquan Wang, Shuying Rao, Junyi Xie, Wanjun Guo, Tao Li, Haiteng Jiang. "SleepVLM: Explainable and Rule-Grounded Sleep Staging via a Vision-Language Model." arXiv preprint, 2026. arXiv:2603.26738 Authors Name Affiliation ORCID Guifeng Deng Zhejiang University 0009-0001-1940-7797 Pan Wang Wenzhou Medical University 0009-0001-6664-6934 Wanjun… See the full description on the dataset page: https://huggingface.co/datasets/Feng613/MASS-EX.texttext-classification10K<n<100K0 likes106 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.