datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.pii-masking-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.pii-masking-400k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.open-pii-masking-500k-ai4privacy
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.pii-masking-200k
Purpose and Features
World's largest open source privacy dataset.
The purpose of the dataset is to train models to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs.
The example texts have 54 PII classes (types of sensitive data), targeting 229 discussion subjects / use cases split across business, education, psychology and legal fields, and 5 interactions styles (e.g. casual conversation, formal document, emails… See the full description on the dataset page: https://huggingface.co/datasets/Isotonic/pii-masking-200k.OpenMath-GSM8K-masked
OpenMath GSM8K Masked
We release a masked version of the GSM8K solutions.
This data can be used to aid synthetic generation of additional solutions for GSM8K dataset
as it is much less likely to lead to inconsistent reasoning compared to using
the original solutions directly.
This dataset was used to construct OpenMathInstruct-1:
a math instruction tuning dataset with 1.8M problem-solution pairs
generated using permissively licensed Mixtral-8x7B model.
For details of how the masked… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMath-GSM8K-masked.OpenMath-MATH-masked
OpenMath GSM8K Masked
We release a masked version of the MATH solutions.
This data can be used to aid synthetic generation of additional solutions for MATH dataset
as it is much less likely to lead to inconsistent reasoning compared to using
the original solutions directly.
This dataset was used to construct OpenMathInstruct-1:
a math instruction tuning dataset with 1.8M problem-solution pairs
generated using permissively licensed Mixtral-8x7B model.
For details of how the masked… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenMath-MATH-masked.pii-masking-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ASR2005Bluesnow/pii-masking-200k.pii-masking-400k
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs.
AI4Privacy Dataset Analytics 📊
Dataset Overview
Total entries: 406,896
Total tokens: 20,564,179
Total PII tokens: 2,357,029
Number of PII classes in public dataset: 17
Number of PII classes in extended dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Ganasekhar/pii-masking-400k.pii-masking-200k
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy. Now it is pii-masking-300k
The purpose of the dataset is to train models to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs.
The example texts have 54 PII classes (types of sensitive data), targeting 229 discussion… See the full description on the dataset page: https://huggingface.co/datasets/saad-kw-almutairi/pii-masking-200k.pii-masking-300k
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs.
Key facts:
OpenPII-220k text entries have 27 PII classes (types of sensitive data), targeting 749 discussion subjects / use cases split across education, health, and psychology. FinPII contains an additional ~20 types tailored to… See the full description on the dataset page: https://huggingface.co/datasets/AdamiTitus/pii-masking-300k.pii-masking-english-1k
Important
This repository contains the English-only subset of the Ai4Privacy PII-Masking-300k Dataset.
The dataset is curated to provide English texts only, while retaining the structure, labeling schema, and licensing of the original dataset.
Licensing
Academic use is encouraged with proper citation provided it follows similar license terms*. Commercial entities should contact us at licensing@ai4privacy.com for licensing inquiries and additional data access.*
Terms… See the full description on the dataset page: https://huggingface.co/datasets/aniket-curlscape/pii-masking-english-1k.pii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/rdany9894/pii-masking-300k.pii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ppalani09/pii-masking-300k.pii-masking-english
Important
This repository contains the English-only subset of the Ai4Privacy PII-Masking-300k Dataset.
The dataset is curated to provide English texts only, while retaining the structure, labeling schema, and licensing of the original dataset.
Licensing
Academic use is encouraged with proper citation provided it follows similar license terms*. Commercial entities should contact us at licensing@ai4privacy.com for licensing inquiries and additional data access.*
Terms… See the full description on the dataset page: https://huggingface.co/datasets/aniket-curlscape/pii-masking-english.pii-masking-300k
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs.
Key facts:
OpenPII-220k text entries have 27 PII classes (types of sensitive data), targeting 749 discussion subjects / use cases split across education, health, and psychology. FinPII contains an additional ~20 types tailored to… See the full description on the dataset page: https://huggingface.co/datasets/saad-kw-almutairi/pii-masking-300k.pii-masking-200k
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy. Now it is pii-masking-300k
The purpose of the dataset is to train models to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs.
The example texts have 54 PII classes (types of sensitive data), targeting 229 discussion… See the full description on the dataset page: https://huggingface.co/datasets/ahczhg/pii-masking-200k.literesearcher-stage2-masked
LiteResearcher Stage-2 (URL-masked, train/val split)
Stage-2 RL data for training a multi-turn search agent, prepared for
verl-style GRPO training with search and
browse tools.
Provenance
Derived from simplex-ai-inc/LiteResearcher-Data
(Apache-2.0). Upstream's pipeline is SFT cold-start → Stage-1 RL → Stage-2 RL;
this repository carries only the Stage-2 portion.
Processing applied on top of upstream:
Stage-2 rows only (upstream Stage 1 is a pure local-RAG warmup… See the full description on the dataset page: https://huggingface.co/datasets/guinansu/literesearcher-stage2-masked.pii-masking-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy. The current flagship is now pii-masking-openpii-1m
The purpose of the dataset is to train models to remove personally identifiable information… See the full description on the dataset page: https://huggingface.co/datasets/shivaniachary123/pii-masking-200k.pii-masking-300k
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs.
Key facts:
OpenPII-220k text entries have 27 PII classes (types of sensitive data), targeting 749 discussion subjects / use cases split across education, health, and psychology. FinPII contains an additional ~20 types tailored to… See the full description on the dataset page: https://huggingface.co/datasets/fuasfgauighsudghaughdoaughsdughdasughoadhg/pii-masking-300k.pii-masking-english-100
Important
This repository contains the English-only subset of the Ai4Privacy PII-Masking-300k Dataset.
The dataset is curated to provide English texts only, while retaining the structure, labeling schema, and licensing of the original dataset.
Licensing
Academic use is encouraged with proper citation provided it follows similar license terms*. Commercial entities should contact us at licensing@ai4privacy.com for licensing inquiries and additional data access.*
Terms… See the full description on the dataset page: https://huggingface.co/datasets/aniket-curlscape/pii-masking-english-100.pii-masking-english-5k
Important
This repository contains the English-only subset of the Ai4Privacy PII-Masking-300k Dataset.
The dataset is curated to provide English texts only, while retaining the structure, labeling schema, and licensing of the original dataset.
Licensing
Academic use is encouraged with proper citation provided it follows similar license terms*. Commercial entities should contact us at licensing@ai4privacy.com for licensing inquiries and additional data access.*
Terms… See the full description on the dataset page: https://huggingface.co/datasets/aniket-curlscape/pii-masking-english-5k.pii_masking_300k_validation_sample_200_english
pii_masking_300k_validation_sample_200_english
PII Masking 300k validation_sample_200.english split
Field
Value
Benchmark
pii_masking_300k
Sub-benchmark
Type
information_extraction
Items
200
Exported from Langfuse.
pii-masking-300k
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs.
Key facts:
OpenPII-220k text entries have 27 PII classes (types of sensitive data), targeting 749 discussion subjects / use cases split across education, health, and psychology. FinPII contains an additional ~20 types tailored to… See the full description on the dataset page: https://huggingface.co/datasets/shivaniachary123/pii-masking-300k.Open-Platypus-Japanese-masked
Open-Platypus-Japanese-masked
LLMの数学能力と推論能力を向上させるために作成したデータセット
garage-bAInd/Open-Platypusをcyberagent/calm3-22b-chatで翻訳
13,881件(13,883件の内2件削除)
ルールベース・機械学習ベースのフィルタリング処理の後、目視確認を行い、個人情報が含まれ得るサンプルについては全件削除対応を実施
データセットの構成とライセンスは以下
Dataset name
Num samples
License type
PRM800k
12298
MIT
SciBench
616
MIT
ARB
713
CC-BY-4.0
Format
{
"idx": インデックス,
"instruction_en": 英語の指示文,
"response_en": 英語の応答文,
"translation_model":… See the full description on the dataset page: https://huggingface.co/datasets/weblab-GENIAC/Open-Platypus-Japanese-masked.pii-masking-400k
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs.
AI4Privacy Dataset Analytics 📊
Dataset Overview
Total entries: 406,896
Total tokens: 20,564,179
Total PII tokens: 2,357,029
Number of PII classes in public dataset: 17
Number of PII classes in extended dataset:… See the full description on the dataset page: https://huggingface.co/datasets/shivaniachary123/pii-masking-400k.aya-ja-evol-instruct-calm3-dpo-masked
aya-ja-evol-instruct-calm3-dpo-masked
LLMの推論能力の向上のためのデータセット
CohereForAI/aya_datasetの日本語パートを抜粋
Evol-instructを適用し、質問を複雑化
進化元の質問文はこのデータセットに含まれていません。team-hatakeyama-phase2/aya-ja-nemotron-dpo-maskedをご参照ください。
Evol-instructの実装はauthor's repositoryに倣って実装しました。
4種のdepth evolvingと1種のbreadth evolving
chosenカラムはcyberagent/calm3-22b-chatで再生成
rejectedカラムは開発時の内部モデルで生成
29,224件(31,295件の内2,071件削除)ルールベース・機械学習ベースのフィルタリング処理の後、目視確認を行い、個人情報が含まれ得るサンプルについては全件削除対応を実施
Format
{… See the full description on the dataset page: https://huggingface.co/datasets/weblab-GENIAC/aya-ja-evol-instruct-calm3-dpo-masked.aya-ja-nemotron-dpo-masked
aya-ja-nemotron-dpo-masked
LLMの推論能力を向上させるためのデータセット
CohereForAI/aya_datasetから日本語パートを抜粋
deepinfraのnvidia/Nemotron-4-340B-Instructで応答を再生成
2024年8月現在ではnvidia/Nemotron-4-340B-Instructは使用不可
5,651件(6,259件の内608件削除)
ルールベース・機械学習ベースのフィルタリング処理の後、目視確認を行い、個人情報が含まれ得るサンプルについては全件削除対応を実施
Format
{
"idx": インデックス,
"prompt": 日本語の指示文,
"chosen": chosenの応答文,
"rejected": rejectedの応答文,
"chosen_model": chosenとしたモデル,
"rejected_model":… See the full description on the dataset page: https://huggingface.co/datasets/weblab-GENIAC/aya-ja-nemotron-dpo-masked.OpenBookQA-Japanese-masked
OpenBookQA-Japanese-masked
与えられた問題に対して4つの選択肢から答えを選択するデータセット
allenai/openbookqaをcyberagent/calm3-22b-chatで翻訳
5,957件
train split: 4,956件(4,957件の内1件削除)
validation split: 500件
test split: 499件(500件の内1件削除)
ルールベース・機械学習ベースのフィルタリング処理の後、目視確認を行い、個人情報が含まれ得るサンプルについては全件削除対応を実施
Format
データセットの構成は以下
{
"idx": ID,
"id": 元ID,
"question_stem_en": 英語の質問文,
"choices_en": {
"text": 選択肢の文章,
"label": 選択肢の記号,
}… See the full description on the dataset page: https://huggingface.co/datasets/weblab-GENIAC/OpenBookQA-Japanese-masked.
