subhash-holla/pii-anon
PII-Anon A CC0 multilingual PII benchmark corpus of 782,677 records carrying 3,107,240 entity annotations across 66 entity types and 60 languages, spanning 7 evaluation dimensions. Each record exposes the five legally-distinct regulatory regime signals (gov-02 / FR-022) as separate reg_* columns — no merged compliance verdict. Train vs. evaluation substrate The 159,891 tier3_evaluation records are the EVALUATION substrate of the 782,677-record corpus… See the full description on the dataset page: https://huggingface.co/datasets/subhash-holla/pii-anon.
PII-Anon
A CC0 multilingual PII benchmark corpus of 782,677 records carrying 3,107,240 entity annotations across 66 entity types and 60 languages, spanning 7 evaluation dimensions. Each record exposes the five legally-distinct regulatory regime signals (gov-02 / FR-022) as separate reg_* columns — no merged compliance verdict.
Train vs. evaluation substrate
The 159,891 tier3_evaluation records are the EVALUATION substrate of the 782,677-record corpus (behavioral-signal / RRS scoring runs on this substrate), NOT the whole corpus.
Synthetic-enrichment disclosure
79.2% of records carry provenance.sourcetype='syntheticlattice_enrichment' (the S-PWR power fill); synthetic power is not external validity.
Statistical power
PII-Anon v2 is powered for all single-factor marginal recall claims (95% Wilson CIs; credential/financial-critical types to ±0.5pp at recall 0.99, standard to ±1pp at 0.98) and for three pre-registered 2-way interactions (language×entity-type on a committed rectangle, domain×track, adversarial-type×entity-type). The corpus carries a committed evaluation lattice powering 17 languages across 11 writing systems (Latin, Han, Japanese, Hangul, Arabic, Devanagari, Cyrillic, Thai, Greek, Bengali, Hebrew) to statistically-calibrated positive-count targets (critical n≥1522, standard n≥753). It is not powered for the full multilingual×entity-type grid or any ≥3-way interaction; those are reported as exploratory. Synthetic-distribution power is not external validity — see the real-data correlation slice.
Baseline Detector Performance
How widely-used PII detectors score on this corpus (test split, language en; 31,048 records), ranked by F2 (recall-weighted — a missed PII is the costly error). Full per-type / per-domain / per-language tables, Wilson CIs, and provenance: baseline_results.json and BASELINES.md.
Power on a committed cell is statistical precision on the SYNTHETIC distribution, NOT external validity; not citable as a standalone recall claim absent the real-data correlation slice (FR-027). Synthetic-only (AX-001).
License
The data is released under CC0-1.0 (public-domain dedication); the accompanying code (loaders, scorers, exporters) is licensed Apache-2.0. The two licenses are distinct — using the data does not subject you to the code license, and vice versa.
