CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai4privacy /pii-masking-openpii-1.5m OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension) 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Overview The OpenPII 1.5M dataset extends OpenPII 1M with a new Asia Pacific corpus, bringing global coverage to 30 languages across Europe, Americas, and Asia Pacific. This is the flagship release of the PII-Masking-3M family, the world's largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.texttoken-classification1M<n<10M21 likes3k downloads4mo agoHugging Face02ai4privacy /pii-masking-openpii-1m OpenPII 1M — Multilingual PII Masking Dataset Overview The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types. Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.texttoken-classification1M<n<10M15 likes2k downloads6mo agoHugging Face03ai4privacy /open-pii-masking-500k-ai4privacy 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.texttext-classification100K<n<1M27 likes1.2k downloads4mo agoHugging Face04ai4privacy /openpii-masking-nano-1k OpenPII Nano: Multilingual PII Masking Sample A nano-sized stratified sample of OpenPII 1.5M, perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every label that exists in the parent dataset is represented in proportion. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format License 1,000 900 100 19 30 37 7… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-nano-1k.texttoken-classification1K<n<10K3 likes111 downloads4mo agoHugging Face05ai4privacy /openpii-masking-micro-100k OpenPII Micro: Multilingual PII Masking Sample A micro-sized stratified sample of OpenPII 1.5M, perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every label that exists in the parent dataset is represented in proportion. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format License 100,000 90,000 10,000 19… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-micro-100k.texttoken-classification100K<n<1M0 likes106 downloads4mo agoHugging Face06Lots-of-LoRAs /task1631_openpi_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1631_openpi_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1631_openpi_answer_generation.texttext-generation1K<n<10K0 likes92 downloads2y agoHugging Face07ai4privacy /openpii-masking-mini-10k OpenPII Masking Mini 10K A compact, stratified subset of ai4privacy/pii-masking-openpii-1m, containing 10,000 samples for rapid experimentation, fine-tuning, and benchmarking of PII detection and masking models. Sampling Methodology Samples were selected using proportional stratified sampling by language: Target count per language = round(lang_proportion × 10,000) — proportional representation. Streaming + reservoir sampling collected 3× the target candidates per… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-mini-10k.texttoken-classification10K<n<100K3 likes75 downloads6mo agoHugging Face08woojin1069 /pii-masking-openpii-1.5m OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension) 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Overview The OpenPII 1.5M dataset extends OpenPII 1M with a new Asia Pacific corpus, bringing global coverage to 30 languages across Europe, Americas, and Asia Pacific. This is the flagship release of the PII-Masking-3M family, the world's largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/woojin1069/pii-masking-openpii-1.5m.token-classification1M<n<10M0 likes74 downloads3mo agoHugging Face09abhi26 /openpipe-dpo-scientific-reasoning Openpipe Dpo Scientific Reasoning This dataset contains 100 high-quality examples for Direct Preference Optimization (DPO) training, formatted for OpenPipe fine-tuning, focused on scientific reasoning and analysis. Dataset Description This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe: OpenAI Chat Format: Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-dpo-scientific-reasoning.texttext-generationn<1K0 likes28 downloads1y agoHugging Face10AdamLucek /open-pii-masking-en-us-30k open-pii-masking-en-us-30k A filtered and transformed subset of ai4privacy/open-pii-masking-500k-ai4privacy filtered for only US English examples where all unique PII labels have been removed and replaced with a [PII] mask. This also includes a new info column with metadata about the count of [PII] masks, and a task column defaulting to 'privacy_masking.' This has resulted in a subset of ~30k examples available within this dataset. For full information about the original… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/open-pii-masking-en-us-30k.texttext-generation10K<n<100K2 likes19 downloads11mo agoHugging Face11abhi26 /openpipe-chat-complete-scientific-reasoning Openpipe Chat Complete Scientific Reasoning This dataset contains 100 high-quality examples for chat completion fine-tuning, formatted for OpenPipe, focused on scientific reasoning and analysis. Dataset Description This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe: OpenAI Chat Format: Standard messages array with… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-chat-complete-scientific-reasoning.texttext-generationn<1K0 likes8 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.