CoolFace
Datasetpublic

Reza2kn/persian-pii-masking-openpii-690k-clean

Persian PII-Masking Combined Corpus, Cleaned Cleaned Persian / Iranian PII-masking token-classification data. This repo combines the audited persona-clean and initial-clean corpora. The combined split contains only rows kept after the source-specific audits. Persona-clean rows: 623890 Initial-clean rows: 224956 Dataset Repo Reza2kn/persian-pii-masking-openpii-690k-clean Schema Rows include: source_text masked_text privacy_mask with label, start… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-pii-masking-openpii-690k-clean.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
1likes20downloads
Dataset Card

Persian PII-Masking Combined Corpus, Cleaned

Cleaned Persian / Iranian PII-masking token-classification data.

This repo combines the audited persona-clean and initial-clean corpora. The combined split contains only rows kept after the source-specific audits.

  • —Persona-clean rows: 623890
  • —Initial-clean rows: 224956

Dataset Repo

Reza2kn/persian-pii-masking-openpii-690k-clean

Schema

Rows include:

  • —source_text
  • —masked_text
  • —privacy_mask with label, start, end, value, label_index
  • —mbert_tokens
  • —mbert_token_classes
  • —language
  • —region
  • —script
  • —uid
  • —split
  • —seed_uid
  • —persona_id (-1 for non-persona initial rows)
  • —rendering_idx
  • —gen_model

Splits

splitrows
train783666
validation32518
test32662
total848846

Cleaning

Cleaning was performed after all rows were embedded with the Mac Studio MLX Qwen3 embedding endpoint. The published clean artifacts remove exact hard failures and keep one representative per full-dataset nearest-neighbor component at cosine >= 0.95.

The local audit manifests used for this repo are uploaded under audit/ where applicable.

Synthetic Data

All personal data is synthetic and generated for PII detection/masking research. Values may be format-valid but do not represent real people.

Attribution

This dataset is localized from the English rows of `ai4privacy/pii-masking-openpii-1m`. License: CC-BY-4.0.