CoolFace
Datasetpublic

tomekkorbak/pile-pii-scrubadub

Dataset Card for pile-pii-scrubadub Dataset Summary This dataset contains text from The Pile, annotated based on the personal idenfitiable information (PII) in each sentence. Each document (row in the dataset) is segmented into sentences, and each sentence is given a score: the percentage of words in it that are classified as PII by Scrubadub. Supported Tasks and Leaderboards [More Information Needed] Languages This dataset is taken… See the full description on the dataset page: https://huggingface.co/datasets/tomekkorbak/pile-pii-scrubadub.

sourceHugging Facemitupdated 4y agoView on Hugging Face
5likes662downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

tomekkorbak/pile-pii-scrubadub · main · files are served by the source, never re-hosted here