CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai4privacy /pii-masking-300k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.texttext-classification100K<n<1M116 likes4.2k downloads4mo agoHugging Face02ai4privacy /pii-masking-200k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.texttext-classification100K<n<1M127 likes3.5k downloads4mo agoHugging Face03ai4privacy /pii-masking-400k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.texttext-classification100K<n<1M63 likes1.3k downloads4mo agoHugging Face04ai4privacy /open-pii-masking-500k-ai4privacy 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.texttext-classification100K<n<1M27 likes1.2k downloads4mo agoHugging Face05Isotonic /pii-masking-200k Purpose and Features World's largest open source privacy dataset. The purpose of the dataset is to train models to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs. The example texts have 54 PII classes (types of sensitive data), targeting 229 discussion subjects / use cases split across business, education, psychology and legal fields, and 5 interactions styles (e.g. casual conversation, formal document, emails… See the full description on the dataset page: https://huggingface.co/datasets/Isotonic/pii-masking-200k.texttext-classification100K<n<1M9 likes323 downloads3y agoHugging Face06nvidia /OpenMath-GSM8K-masked OpenMath GSM8K Masked We release a masked version of the GSM8K solutions. This data can be used to aid synthetic generation of additional solutions for GSM8K dataset as it is much less likely to lead to inconsistent reasoning compared to using the original solutions directly. This dataset was used to construct OpenMathInstruct-1: a math instruction tuning dataset with 1.8M problem-solution pairs generated using permissively licensed Mixtral-8x7B model. For details of how the masked… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMath-GSM8K-masked.textquestion-answering1K<n<10K12 likes198 downloads3y agoHugging Face07nvidia /OpenMath-MATH-masked OpenMath GSM8K Masked We release a masked version of the MATH solutions. This data can be used to aid synthetic generation of additional solutions for MATH dataset as it is much less likely to lead to inconsistent reasoning compared to using the original solutions directly. This dataset was used to construct OpenMathInstruct-1: a math instruction tuning dataset with 1.8M problem-solution pairs generated using permissively licensed Mixtral-8x7B model. For details of how the masked… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMath-MATH-masked.textquestion-answering1K<n<10K9 likes150 downloads3y agoHugging Face08ASR2005Bluesnow /pii-masking-200k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ASR2005Bluesnow/pii-masking-200k.texttext-classification100K<n<1M0 likes78 downloads23d agoHugging Face09Ganasekhar /pii-masking-400k Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. AI4Privacy Dataset Analytics 📊 Dataset Overview Total entries: 406,896 Total tokens: 20,564,179 Total PII tokens: 2,357,029 Number of PII classes in public dataset: 17 Number of PII classes in extended dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Ganasekhar/pii-masking-400k.texttext-classification100K<n<1M0 likes55 downloads7mo agoHugging Face10saad-kw-almutairi /pii-masking-200k Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy. Now it is pii-masking-300k The purpose of the dataset is to train models to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs. The example texts have 54 PII classes (types of sensitive data), targeting 229 discussion… See the full description on the dataset page: https://huggingface.co/datasets/saad-kw-almutairi/pii-masking-200k.texttext-classification100K<n<1M0 likes50 downloads9mo agoHugging Face11AdamiTitus /pii-masking-300k Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. Key facts: OpenPII-220k text entries have 27 PII classes (types of sensitive data), targeting 749 discussion subjects / use cases split across education, health, and psychology. FinPII contains an additional ~20 types tailored to… See the full description on the dataset page: https://huggingface.co/datasets/AdamiTitus/pii-masking-300k.texttext-classification100K<n<1M1 likes50 downloads8mo agoHugging Face12aniket-curlscape /pii-masking-english-1k Important This repository contains the English-only subset of the Ai4Privacy PII-Masking-300k Dataset. The dataset is curated to provide English texts only, while retaining the structure, labeling schema, and licensing of the original dataset. Licensing Academic use is encouraged with proper citation provided it follows similar license terms*. Commercial entities should contact us at licensing@ai4privacy.com for licensing inquiries and additional data access.* Terms… See the full description on the dataset page: https://huggingface.co/datasets/aniket-curlscape/pii-masking-english-1k.texttext-classification1K<n<10K0 likes45 downloads1y agoHugging Face13rdany9894 /pii-masking-300k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/rdany9894/pii-masking-300k.texttext-classification100K<n<1M0 likes43 downloads3d agoHugging Face14ppalani09 /pii-masking-300k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ppalani09/pii-masking-300k.texttext-classification100K<n<1M0 likes41 downloads3mo agoHugging Face15aniket-curlscape /pii-masking-english Important This repository contains the English-only subset of the Ai4Privacy PII-Masking-300k Dataset. The dataset is curated to provide English texts only, while retaining the structure, labeling schema, and licensing of the original dataset. Licensing Academic use is encouraged with proper citation provided it follows similar license terms*. Commercial entities should contact us at licensing@ai4privacy.com for licensing inquiries and additional data access.* Terms… See the full description on the dataset page: https://huggingface.co/datasets/aniket-curlscape/pii-masking-english.texttext-classification10K<n<100K0 likes33 downloads1y agoHugging Face16saad-kw-almutairi /pii-masking-300k Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. Key facts: OpenPII-220k text entries have 27 PII classes (types of sensitive data), targeting 749 discussion subjects / use cases split across education, health, and psychology. FinPII contains an additional ~20 types tailored to… See the full description on the dataset page: https://huggingface.co/datasets/saad-kw-almutairi/pii-masking-300k.texttext-classification100K<n<1M0 likes33 downloads9mo agoHugging Face17ahczhg /pii-masking-200k Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy. Now it is pii-masking-300k The purpose of the dataset is to train models to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs. The example texts have 54 PII classes (types of sensitive data), targeting 229 discussion… See the full description on the dataset page: https://huggingface.co/datasets/ahczhg/pii-masking-200k.texttext-classification100K<n<1M0 likes33 downloads7mo agoHugging Face18guinansu /literesearcher-stage2-masked LiteResearcher Stage-2 (URL-masked, train/val split) Stage-2 RL data for training a multi-turn search agent, prepared for verl-style GRPO training with search and browse tools. Provenance Derived from simplex-ai-inc/LiteResearcher-Data (Apache-2.0). Upstream's pipeline is SFT cold-start → Stage-1 RL → Stage-2 RL; this repository carries only the Stage-2 portion. Processing applied on top of upstream: Stage-2 rows only (upstream Stage 1 is a pure local-RAG warmup… See the full description on the dataset page: https://huggingface.co/datasets/guinansu/literesearcher-stage2-masked.textquestion-answering10K<n<100K0 likes31 downloads1mo agoHugging Face19shivaniachary123 /pii-masking-200k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy. The current flagship is now pii-masking-openpii-1m The purpose of the dataset is to train models to remove personally identifiable information… See the full description on the dataset page: https://huggingface.co/datasets/shivaniachary123/pii-masking-200k.texttext-classification100K<n<1M0 likes26 downloads4mo agoHugging Face20fuasfgauighsudghaughdoaughsdughdasughoadhg /pii-masking-300k Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. Key facts: OpenPII-220k text entries have 27 PII classes (types of sensitive data), targeting 749 discussion subjects / use cases split across education, health, and psychology. FinPII contains an additional ~20 types tailored to… See the full description on the dataset page: https://huggingface.co/datasets/fuasfgauighsudghaughdoaughsdughdasughoadhg/pii-masking-300k.texttext-classification100K<n<1M0 likes25 downloads5mo agoHugging Face21aniket-curlscape /pii-masking-english-100 Important This repository contains the English-only subset of the Ai4Privacy PII-Masking-300k Dataset. The dataset is curated to provide English texts only, while retaining the structure, labeling schema, and licensing of the original dataset. Licensing Academic use is encouraged with proper citation provided it follows similar license terms*. Commercial entities should contact us at licensing@ai4privacy.com for licensing inquiries and additional data access.* Terms… See the full description on the dataset page: https://huggingface.co/datasets/aniket-curlscape/pii-masking-english-100.texttext-classificationn<1K0 likes24 downloads1y agoHugging Face22aniket-curlscape /pii-masking-english-5k Important This repository contains the English-only subset of the Ai4Privacy PII-Masking-300k Dataset. The dataset is curated to provide English texts only, while retaining the structure, labeling schema, and licensing of the original dataset. Licensing Academic use is encouraged with proper citation provided it follows similar license terms*. Commercial entities should contact us at licensing@ai4privacy.com for licensing inquiries and additional data access.* Terms… See the full description on the dataset page: https://huggingface.co/datasets/aniket-curlscape/pii-masking-english-5k.texttext-classification1K<n<10K0 likes19 downloads1y agoHugging Face23orgrctera /pii_masking_300k_validation_sample_200_english pii_masking_300k_validation_sample_200_english PII Masking 300k validation_sample_200.english split Field Value Benchmark pii_masking_300k Sub-benchmark Type information_extraction Items 200 Exported from Langfuse. textquestion-answeringn<1K0 likes17 downloads7mo agoHugging Face24shivaniachary123 /pii-masking-300k Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. Key facts: OpenPII-220k text entries have 27 PII classes (types of sensitive data), targeting 749 discussion subjects / use cases split across education, health, and psychology. FinPII contains an additional ~20 types tailored to… See the full description on the dataset page: https://huggingface.co/datasets/shivaniachary123/pii-masking-300k.texttext-classification100K<n<1M0 likes14 downloads5mo agoHugging Face25weblab-GENIAC /Open-Platypus-Japanese-maskedgated Open-Platypus-Japanese-masked LLMの数学能力と推論能力を向上させるために作成したデータセット garage-bAInd/Open-Platypusをcyberagent/calm3-22b-chatで翻訳 13,881件(13,883件の内2件削除) ルールベース・機械学習ベースのフィルタリング処理の後、目視確認を行い、個人情報が含まれ得るサンプルについては全件削除対応を実施 データセットの構成とライセンスは以下 Dataset name Num samples License type PRM800k 12298 MIT SciBench 616 MIT ARB 713 CC-BY-4.0 Format { "idx": インデックス, "instruction_en": 英語の指示文, "response_en": 英語の応答文, "translation_model":… See the full description on the dataset page: https://huggingface.co/datasets/weblab-GENIAC/Open-Platypus-Japanese-masked.texttext-generation10K<n<100K1 likes13 downloads2y agoHugging Face26shivaniachary123 /pii-masking-400k Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. AI4Privacy Dataset Analytics 📊 Dataset Overview Total entries: 406,896 Total tokens: 20,564,179 Total PII tokens: 2,357,029 Number of PII classes in public dataset: 17 Number of PII classes in extended dataset:… See the full description on the dataset page: https://huggingface.co/datasets/shivaniachary123/pii-masking-400k.texttext-classification100K<n<1M0 likes12 downloads5mo agoHugging Face27weblab-GENIAC /aya-ja-evol-instruct-calm3-dpo-maskedgated aya-ja-evol-instruct-calm3-dpo-masked LLMの推論能力の向上のためのデータセット CohereForAI/aya_datasetの日本語パートを抜粋 Evol-instructを適用し、質問を複雑化 進化元の質問文はこのデータセットに含まれていません。team-hatakeyama-phase2/aya-ja-nemotron-dpo-maskedをご参照ください。 Evol-instructの実装はauthor's repositoryに倣って実装しました。 4種のdepth evolvingと1種のbreadth evolving chosenカラムはcyberagent/calm3-22b-chatで再生成 rejectedカラムは開発時の内部モデルで生成 29,224件(31,295件の内2,071件削除)ルールベース・機械学習ベースのフィルタリング処理の後、目視確認を行い、個人情報が含まれ得るサンプルについては全件削除対応を実施 Format {… See the full description on the dataset page: https://huggingface.co/datasets/weblab-GENIAC/aya-ja-evol-instruct-calm3-dpo-masked.texttext-generation10K<n<100K8 likes10 downloads2y agoHugging Face28weblab-GENIAC /aya-ja-nemotron-dpo-maskedgated aya-ja-nemotron-dpo-masked LLMの推論能力を向上させるためのデータセット CohereForAI/aya_datasetから日本語パートを抜粋 deepinfraのnvidia/Nemotron-4-340B-Instructで応答を再生成 2024年8月現在ではnvidia/Nemotron-4-340B-Instructは使用不可 5,651件(6,259件の内608件削除) ルールベース・機械学習ベースのフィルタリング処理の後、目視確認を行い、個人情報が含まれ得るサンプルについては全件削除対応を実施 Format { "idx": インデックス, "prompt": 日本語の指示文, "chosen": chosenの応答文, "rejected": rejectedの応答文, "chosen_model": chosenとしたモデル, "rejected_model":… See the full description on the dataset page: https://huggingface.co/datasets/weblab-GENIAC/aya-ja-nemotron-dpo-masked.texttext-generation1K<n<10K4 likes7 downloads2y agoHugging Face29weblab-GENIAC /OpenBookQA-Japanese-maskedgated OpenBookQA-Japanese-masked 与えられた問題に対して4つの選択肢から答えを選択するデータセット allenai/openbookqaをcyberagent/calm3-22b-chatで翻訳 5,957件 train split: 4,956件(4,957件の内1件削除) validation split: 500件 test split: 499件(500件の内1件削除) ルールベース・機械学習ベースのフィルタリング処理の後、目視確認を行い、個人情報が含まれ得るサンプルについては全件削除対応を実施 Format データセットの構成は以下 { "idx": ID, "id": 元ID, "question_stem_en": 英語の質問文, "choices_en": { "text": 選択肢の文章, "label": 選択肢の記号, }… See the full description on the dataset page: https://huggingface.co/datasets/weblab-GENIAC/OpenBookQA-Japanese-masked.tabulartext-generation1K<n<10K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.