Ganasekhar/pii-masking-400k
Purpose and Features π World's largest open dataset for privacy masking π The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. AI4Privacy Dataset Analytics π Dataset Overview Total entries: 406,896 Total tokens: 20,564,179 Total PII tokens: 2,357,029 Number of PII classes in public dataset: 17 Number of PII classes inβ¦ See the full description on the dataset page: https://huggingface.co/datasets/Ganasekhar/pii-masking-400k.
056
Duplicate from ai4privacy/pii-masking-400k
