CoolFace
Datasetpublic

mapo80/aliasit-pii-dataset-v4-small

aliasit-pii-dataset-v4-small One fifth of the training split of aliasit-pii-dataset-v4, sampled uniformly on ids so the category distribution survives, with validation and test kept whole. For training runs that fail fast and still evaluate on the real thing. What this is An Italian-only PII dataset in text + character span form. Spans are character offsets into the raw text, end is exclusive, and they are independent of any tokenizer: converting them to token… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v4-small.

sourceHugging Faceotherupdated 25d agoView on Hugging Face
0likes61downloads
Dataset Card

aliasit-pii-dataset-v4-small

One fifth of the training split of aliasit-pii-dataset-v4, sampled uniformly on ids so the category distribution survives, with validation and test kept whole. For training runs that fail fast and still evaluate on the real thing.

What this is

An Italian-only PII dataset in text + character span form. Spans are character offsets into the raw text, end is exclusive, and they are independent of any tokenizer: converting them to token labels is the consumer's job, so the dataset does not bake in a tokenization choice.

LanguageItalian (it)
Categories79 in the taxonomy, 42 used here
Documents49,349
Entities919,352
Taxonomy version3.0.0
Dataset revision1.0.0
SplitDocumentsEntitiesTraining
train19,930306,058allowed
validation14,776308,483allowed
test14,643304,811forbidden

test must not be used for training, threshold calibration, model selection, distillation or hyper-parameter tuning. An evaluation set that cannot fail is not an evaluation set: use validation for anything that adapts to the data.

Licensing — read this before redistributing

This dataset is derived from many sources with different licences. Most are CC-BY-4.0, MIT or Apache-2.0 and are redistributed under their own terms, with attribution listed below.

2 of the sources declare no licence upstream, and they account for 3,794 documents (7.7% of this dataset):

SourceDocumentsShare
DeepMount00/pii-masking-ita2,2634.59%
DeepMount00/GLINER_ITA1,5313.10%

Without a known upstream licence there is no established right to redistribute the material derived from them. They were included on an explicit decision by the dataset author, who accepted that risk. The consequence for you: this dataset cannot assert a single licence over its whole content, the other licence tag reflects exactly that, and if you redistribute it further the same uncertainty travels with you. If your use case needs a clean chain of title, drop those sources — they are identified per record in the prov_dataset column.

Sources and attribution

SourceLicenceAttribution requiredDocumentsPin
rizzoaiacademy/anonimizzazione-testi-italiano-cleanMITno27,75150163bcda973efe8…
klusai/ds-kp-general-it-50kCC-BY-4.0yes8,5871f46211fd5ba203c…
ai4privacy/pii-masking-openpii-1mCC-BY-4.0no2,842ecfdc547f4a09556…
DeepMount00/pii-masking-itaUNKNOWNno2,263
src/genera_italiano.pyGENERATED IN-REPOno1,692see the code
DeepMount00/GLINER_ITAUNKNOWNno1,531b56bbd496b9d77bf…
gretelai/synthetic_pii_finance_multilingualAPACHE-2.0no1,1717b844d16738527a0…
urchade/synthetic-pii-ner-mistral-v1APACHE-2.0no1,109392d6853aaec2807…
dr3x1/rizzo-pii-security-itMITno658125fc0865df7325f…
Ethosoft/TR-DocVQA-SynthCC-BY-4.0yes3861f1ee4ba172c85a4…
nvidia/Nemotron-PIICC-BY-4.0yes243b70ffaf5ff39e079…
Ministero della Salute, Manuale ICD-9-CM versione italiana 2007 (IODL 2.0)IODL-2.0yes221sha256 0550178d66754a7f…
albertobarnabo/synthetic-receipts-ocrAPACHE-2.0no172ed46e02b9b1136f6…
dossier-legal/italian-legal-corpusCC-BY-4.0yes106e503a93f124d76b2…
Ministero dell'Interno - ANPR, archivio dei comuni (CC-BY-4.0)CC-BY-4.0yes89sha256 2a83f63414e43eae…
kurkowski/synthetic-contextual-anonymizer-datasetMITno844b3170d51764950f…
E3-JSI/synthetic-multi-pii-ner-v1MITno832fe1b17ea2d82d80…
mik3ml/personas-italianAPACHE-2.0no7237dffd4490e50992…
istat-ai/court-rulings-coiAPACHE-2.0no650c98c00bb33c3ddc…
DM 23 dicembre 1976, implementato in src/codice_fiscale.pyGENERATED IN-REPOno61see the code
IVASS - Istituto per la Vigilanza sulle Assicurazioni, Lista imprese assicurative vigilate (CC-BY-4.0)CC-BY-4.0yes59sha256 de199f2fc46ae4bf…
Ministero dell'Interno - ANPR, tabella 2 stati esteri (CC-BY-4.0)CC-BY-4.0yes49sha256 9a1843f89ad6a0f6…
Agenzia Italiana del Farmaco (AIFA), liste dei farmaci di classe A (CC-BY-4.0)CC-BY-4.0yes28sha256 216a36771053e1ea…
Toridion/lindisfarne-m1CC-BY-4.0yes24962c818e10a6e329…
huseyinatahaninan/ContextualIntegritySyntheticDatasetAPACHE-2.0no2e6fe38ea25684131…
DataDock/geonamesCC-BY-4.0yes1994823ca4a8b80a2…

The sources marked yes are redistributed under licences that require attribution. If you use this dataset, carry the following credits:

  • Agenzia Italiana del Farmaco (AIFA), liste dei farmaci di classe A (CC-BY-4.0)
  • Ethosoft, TR-DocVQA-Synth (CC-BY-4.0)
  • GeoNames (CC-BY-4.0), mirror DataDock/geonames
  • IVASS - Istituto per la Vigilanza sulle Assicurazioni, Lista imprese assicurative vigilate (CC-BY-4.0)
  • KlusAI, ds-kp-general-it-50k (CC-BY-4.0)
  • Ministero dell'Interno - ANPR, archivio dei comuni (CC-BY-4.0)
  • Ministero dell'Interno - ANPR, tabella 2 stati esteri (CC-BY-4.0)
  • Ministero della Salute, Manuale ICD-9-CM versione italiana 2007 (IODL 2.0)
  • NVIDIA, Nemotron-PII (CC-BY-4.0)
  • Toridion, Lindisfarne M1 (CC-BY-4.0)
  • dossier-legal, italian-legal-corpus (CC-BY-4.0)

What changed in this revision

Some categories were merged into their parent because a model cannot choose between them: the text carries no signal that separates the two, and the parent has orders of magnitude more examples. Measured on a gold set of real Italian documents, the model trained on the previous revision wrote date on 22 dates of birth out of 22 — never once the subcategory.

merged awayintowhy
birth_placecitysame surface form, city has far more support
date_of_birthdatesame surface form, date has far more support
landline_phone_numberphone_numbersame surface form, phone_number has far more support
mobile_phone_numberphone_numbersame surface form, phone_number has far more support

For anonymisation nothing is lost: a date of birth masked as date is still masked. What is lost is the distinction, and if you need it you have to recover it from the sources.

Composite address spans were decomposed into their components (street_name, building_number, postal_code, city, italian_province) wherever that was possible, so that the same text pattern always carries the same annotation. The decomposition was proposed by rizzoaiacademy/rizzo-pii-0.3B, a model that has no address category: it is another model's annotation, not ground truth. 96.1% of address spans were decomposed; the rest stay whole rather than risk an invented component.

Personal data

Every source declares its content synthetic. None of them documents a formal verification that no real personal data slipped in, so this dataset inherits that level of assurance and not a higher one: it is a documented assumption, not a fact we verified. Treat it as synthetic training material, not as proof that no real person is described.

Format

id                  stable document id, shared with the source dataset
text                raw document text, never modified
language            always 'it'
split               train | validation | test | benchmark
ent_start[]         character offset of each entity, inclusive
ent_len[]           length in characters; end = start + len, exclusive
ent_label[]         one of the 44 categories
ent_source_label[]  the label the source used, before mapping
ent_source_dataset[] which source produced that annotation
unm_*[]             spans the source annotated and we did NOT map, with the reason
prov_*              dataset, repo_id, revision, original_id, licence, register, ...
fp_*                fingerprints: exact, normalized, entity_set, template
cont_*              contamination flags for the two reference models

*`unm_` matters more than it looks.* A span listed there is PII that the source annotated and this dataset deliberately does not: the category is out of scope, or the source filled it with values that are not what the category means. It is not* an oversight, and it is not a negative example either.

A document may contain strings that are PII of an excluded category and carry no annotation for them. Here not annotated does not mean not present.

Splits and leakage

Splits are built so that a template family never crosses them: the skeleton of a document, with entity values masked and digits normalised, belongs to exactly one split. The build verifies this and stops if it does not hold, so a model evaluated here cannot be scoring for remembering a layout it saw in training.

This is the small variant: the training split is a uniform sample on document ids, which preserves the category distribution of the full dataset, while validation and test are kept whole — the rarest categories there sit exactly at the readability threshold, and sampling them would push their per-category F1 below the point where it means anything. Ids match the full dataset, so a record here is findable there.

Checksums

FileSHA-256Bytes
aliasit-pii-dataset-v4-small.test.parquet67b82e73ed10d8c2…4,879,050
aliasit-pii-dataset-v4-small.train.parquet4f97e92484b58f59…5,802,731
aliasit-pii-dataset-v4-small.validation.parquetbeb0967528615692…4,920,079

manifest.json ships with the data. Verify it before you trust that what you have is what produced any number quoted here — if a hash differs, this card does not describe your copy.

Citation

bibtex
@dataset{aliasit_pii_aliasit_pii_dataset_v4_small,
  title   = {aliasit-pii-dataset-v4-small},
  author  = {Polito, Matteo},
  year    = {2026},
  note    = {Italian PII dataset, 44 entity types, revision 1.0.0},
  url     = {https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v4-small}
}

Derived from the sources listed above, under their respective licences, with the caveat stated in the licensing section.