CoolFace
Datasetpublic

xorushi/roberta-pii-synth

Synthetic PII Detection Dataset (RoBERTa-PII-Synth) A large-scale, fully synthetic dataset for training token-classification models to detect Personally Identifiable Information (PII) in realistic text. This dataset was built using an enhanced synthetic generation pipeline, designed to better capture the linguistic and formatting variability of real-world user text. All samples are fully artificial — no real people or identifiers appear anywhere. 📘 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/roberta-pii-synth.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes87downloads
settings

This repository belongs to xorushi on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameroberta-pii-synth
visibilitypublic
licencemit
gatedno
ownerxorushi
Account settings
xorushi/roberta-pii-synth · CoolFace