mapo80/aliasit-pii-dataset-v4-small
aliasit-pii-dataset-v4-small One fifth of the training split of aliasit-pii-dataset-v4, sampled uniformly on ids so the category distribution survives, with validation and test kept whole. For training runs that fail fast and still evaluate on the real thing. What this is An Italian-only PII dataset in text + character span form. Spans are character offsets into the raw text, end is exclusive, and they are independent of any tokenizer: converting them to token… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v4-small.
065
aliasit-pii: aliasit-pii-dataset-v4-small
aliasit-pii: aliasit-pii-dataset-v4-small
initial commit
