orgrctera/pii_masking_300k_information_extraction
PII Masking 300k — Information Extraction Dataset summary This repository hosts a validation sample of the PII Masking 300k benchmark for the information extraction track: models must identify personally identifiable information (PII) in text and produce structured extractions (slot-filling JSON), optional token-level BIO labels, and span-based annotations for masking or redaction workflows. The full PII Masking 300k suite is designed to stress-test… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/pii_masking_300k_information_extraction.
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Push pii_masking_300k_information_extraction (200 items, 1 splits) from Langfuse
Add dataset card
Push pii_masking_300k_information_extraction (200 items, 1 splits) from Langfuse
initial commit
