575-lab/kiji-inspector-reviewed-pairs
Kiji PII Detection Training Data Synthetic multilingual dataset for training PII (Personally Identifiable Information) detection models with token-level entity annotations and coreference resolution. Dataset Summary Samples 99,990 (train: 89,991, test: 9,999) Languages 6 (Dutch, Spanish, German, English, Danish, French) Countries 20 PII entity types 26 Total entity annotations 814,306 (avg 8.1 per sample) Coreference clusters 142,142 (99% of… See the full description on the dataset page: https://huggingface.co/datasets/575-lab/kiji-inspector-reviewed-pairs.
033
Upload README.md with huggingface_hub
Add audit ledger
Upload dataset
initial commit
