TonyYun/pii-shield-benchmark
PII Shield Benchmark Five public PII-detection datasets rewritten into one record shape, so models can be trained and evaluated across all of them without writing five parsers. 2,301,892 records carrying 13,338,839 labeled spans in 25 European languages. Nothing here is new text. Every document comes unchanged from one of the five source datasets below; this repository harmonizes the containers — one schema, one label taxonomy, one file format (Parquet, zstd) — and publishes a… See the full description on the dataset page: https://huggingface.co/datasets/TonyYun/pii-shield-benchmark.
v2: cleaned annotations — card
v2: cleaned annotations — manifest
v2: cleaned annotations — openpii-1m
v2: cleaned annotations — privy
v2: cleaned annotations — nemotron-pii
v2: cleaned annotations — mapa-testdocs
v2: cleaned annotations — gretel-pii-en-v1
Card: remove the flags documentation (field removed from the data)
Manifest for the flag-less rebuild
Rebuild mapa-testdocs without the per-span flags field
Rebuild mapa-sentences without the per-span flags field
Rebuild privy without the per-span flags field
Rebuild gretel-pii-en-v1 without the per-span flags field
Rebuild nemotron-pii without the per-span flags field
Rebuild openpii-1m without the per-span flags field
card: document the non_literal_value flag
manifest: non_literal_value flag counts
nemotron: flag 563 non-literal email spans (non_literal_value)
remove attribution line
unified data: privy
unified data: openpii-1m
unified data: nemotron-pii
unified data: mapa-testdocs
unified data: mapa-sentences
unified data: gretel-pii-en-v1
build manifest
dataset card
initial commit
