Reza2kn/persian-pii-masking-openpii-690k-clean
Persian PII-Masking Combined Corpus, Cleaned Cleaned Persian / Iranian PII-masking token-classification data. This repo combines the audited persona-clean and initial-clean corpora. The combined split contains only rows kept after the source-specific audits. Persona-clean rows: 623890 Initial-clean rows: 224956 Dataset Repo Reza2kn/persian-pii-masking-openpii-690k-clean Schema Rows include: source_text masked_text privacy_mask with label, start… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-pii-masking-openpii-690k-clean.
Persian PII-Masking Combined Corpus, Cleaned
Cleaned Persian / Iranian PII-masking token-classification data.
This repo combines the audited persona-clean and initial-clean corpora. The combined split contains only rows kept after the source-specific audits.
- Persona-clean rows:
623890 - Initial-clean rows:
224956
Dataset Repo
Reza2kn/persian-pii-masking-openpii-690k-clean
Schema
Rows include:
source_textmasked_textprivacy_maskwithlabel,start,end,value,label_indexmbert_tokensmbert_token_classeslanguageregionscriptuidsplitseed_uidpersona_id(-1for non-persona initial rows)rendering_idxgen_model
Splits
Cleaning
Cleaning was performed after all rows were embedded with the Mac Studio MLX Qwen3 embedding endpoint. The published clean artifacts remove exact hard failures and keep one representative per full-dataset nearest-neighbor component at cosine >= 0.95.
The local audit manifests used for this repo are uploaded under audit/ where applicable.
Synthetic Data
All personal data is synthetic and generated for PII detection/masking research. Values may be format-valid but do not represent real people.
Attribution
This dataset is localized from the English rows of `ai4privacy/pii-masking-openpii-1m`. License: CC-BY-4.0.
