CoolFace
Datasetpublic

mapo80/aliasit-pii-dataset-v4-small

aliasit-pii-dataset-v4-small One fifth of the training split of aliasit-pii-dataset-v4, sampled uniformly on ids so the category distribution survives, with validation and test kept whole. For training runs that fail fast and still evaluate on the real thing. What this is An Italian-only PII dataset in text + character span form. Spans are character offsets into the raw text, end is exclusive, and they are independent of any tokenizer: converting them to token… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v4-small.

sourceHugging Faceotherupdated 27d agoView on Hugging Face
0likes65downloads
3 commits on main
2c23c3a27d ago

aliasit-pii: aliasit-pii-dataset-v4-small

mapo80
4f6b2e927d ago

aliasit-pii: aliasit-pii-dataset-v4-small

mapo80
d83474f27d ago

initial commit

mapo80