CoolFace
Datasetpublic

mapo80/aliasit-pii-dataset-v5

aliasit-pii-dataset-v5 Italian PII dataset in canonical text + character span form, the 42 entity types of v4, with the form variety and the document length that a gold set of real Italian documents showed were missing: perturbed surface forms, composed multi-section documents with anaphoric surname references, and numeric material deliberately labelled O. Read this first: the text is rewritten Every dataset in this family before v5 kept the source text byte for… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v5.

sourceHugging Faceotherupdated 24d agoView on Hugging Face
0likes54downloads
Dataset Card

aliasit-pii-dataset-v5

Italian PII dataset in canonical text + character span form, the 42 entity types of v4, with the form variety and the document length that a gold set of real Italian documents showed were missing: perturbed surface forms, composed multi-section documents with anaphoric surname references, and numeric material deliberately labelled O.

Read this first: the text is rewritten

Every dataset in this family before v5 kept the source text byte for byte. This one does not. The entity values still come from the sources, but their surface form is rewritten on purpose, and documents are recomposed into longer ones. If you need the original wording, use `aliasit-pii-dataset-v4`, which is frozen and unchanged.

Why: a model trained on v4 scores 1.000 exact on 14,087 spans of unseen documents from one synthetic source, and 0.447 on a gold set of real Italian documents. It had learned the grammar of the generators, not Italian. The clearest single symptom: 07118 is recognised as a postal code in Via Trento 44 07118 Sassari and missed in Viale Manzoni, 20 - 07118 NOVARA.

What this is

LanguageItalian (it)
Categories79 in the taxonomy, 42 used here
Documents54,734
Entities2,563,450
Taxonomy version3.0.0
Dataset revision1.0.0
Derived fromaliasit-pii-dataset-v4
SplitDocumentsEntitiesTraining
train37,9981,604,108allowed
validation5,584321,072allowed
test5,568317,198forbidden
validation_perturbata5,584321,072allowed

`validation_perturbata` is not an independent split. It is the same validation documents rewritten with every form perturbation at full rate and a different seed. It ships because it is the set on which the training recipe for this dataset was selected, and publishing it is the only way that selection can be redone from outside. Do not add it to validation and call the sum an evaluation set: you would be counting the same documents twice.

There is no `benchmark` split. aliasit-pii-benchmark-v1 is frozen elsewhere and is not touched by this dataset.

How the text was rewritten

Four transformations, all deterministic given a seed, all declared in configs/dataset.yaml before the first build. The proportions below are the configuration; the distributions further down are what was measured on the built dataset by re-reading the text.

whatrateseed
Surface formsprefixes of codice_fiscale / partita_iva, address punctuation and postal-code position, date formats, capitalisation of names and organisations60% of ids, 70% of addresses, 60% of dates, 55% of persons, 45% of organisations, 8% of documents fully upper-cased20260828
Compositiondocuments are consumed into multi-section ones with a real institution in the header, numbered sections, and a closing part where names come back as bare surnames at least 1,500 characters after the full name70% of documents, target 6,500 characters20260828
Hard negativesstatutory references, protocol numbers, amounts, page and resolution numbers — numeric material deliberately left O6 per composed document, 35% of the rest get 320260828
Institution namesreal Italian public-body names in headers, from a static committed list with declared provenance1 per composed document, 45% get a second20260828

Nothing numeric that is written stays unannotated by accident. The hard negatives are exactly the categories that the annotation guidelines for the gold set exclude by name — statutory citations, page numbers, amounts (monetary_amount is out of the taxonomy), and act numbers not tied to a person. No negative sentence names anybody, and no negative sentence contains a standalone date: writing a date and not annotating it would teach a model that dates are not personal data, which is the worst thing an anonymisation dataset can do.

What that produced, measured on the built dataset

Categorysurface forms above 5%distribution
codice_fiscale5codice fiscale 31%, cf 24%, inline 18%, c.f. 15%, cod. fisc. 12%
partita_iva5p.iva 29%, partita iva 27%, inline 19%, piva 15%, p. iva 11%
date5gg/mm/aaaa 39%, gg mese aaaa 26%, gg-mm-aaaa 10%, iso 9%, gg.mm.aaaa 9%
person4titolo 40%, una_parola 25%, maiuscolo 24%, cognome_caps 11%
organization4titolo 55%, maiuscolo 22%, discorsivo 15%, sigla 6%
indirizzo6virgola 41%, piano 17%, n_civico 12%, virgola_trattino 11%, cap_dopo_citta 10%, a_capo 9%
v4 trainingv5 training
Document length, p901,937 chars8,589 chars
Longest document2,992 chars11,065 chars
person that are a bare surname resumed from a full name 1,500+ chars earlier18.3%
Documents with 3+ numeric sequences labelled O9.7%53.5%

What is NOT in here

This dataset contains no real personal data collected from the web. The project also holds a private gold set of real Italian administrative and judicial documents, and a larger unannotated collection from the same sources. Neither is published, in any form — not the texts, not extracts, not annotations.

That is checked, not asserted. Every document of this dataset is compared against every one of those real documents at four levels — identical text, identical after normalising form, same template family, and shingle containment that survives a rewrite. Result: 0 collisions across 54,734 documents.

Licensing — read this before redistributing

Inherited unchanged from v4: this dataset adds no source, it reuses them. Most are CC-BY-4.0, MIT or Apache-2.0 and are redistributed under their own terms, with attribution listed below.

2 of the sources declare no licence upstream, and they account for 3,261 documents (6.0% of this dataset):

SourceDocumentsShare
DeepMount00/pii-masking-ita1,9253.52%
DeepMount00/GLINER_ITA1,3362.44%

Without a known upstream licence there is no established right to redistribute the material derived from them. They were included on an explicit decision by the dataset author, who accepted that risk. The consequence for you: this dataset cannot assert a single licence over its whole content, the other licence tag reflects exactly that, and if you redistribute it further the same uncertainty travels with you. If your use case needs a clean chain of title, drop those sources — they are identified per record in the prov_dataset column.

Sources and attribution

SourceLicenceAttribution requiredDocumentsPin
rizzoaiacademy/anonimizzazione-testi-italiano-cleanMITno21,84350163bcda973efe8…
aliasit_composto_it (documents composed in-repo from the sources above)DERIVEDno11,277see the code
klusai/ds-kp-general-it-50kCC-BY-4.0yes9,4921f46211fd5ba203c…
ai4privacy/pii-masking-openpii-1mCC-BY-4.0no3,166ecfdc547f4a09556…
DeepMount00/pii-masking-itaUNKNOWNno1,925
DeepMount00/GLINER_ITAUNKNOWNno1,336b56bbd496b9d77bf…
gretelai/synthetic_pii_finance_multilingualAPACHE-2.0no1,2787b844d16738527a0…
src/genera_italiano.pyGENERATED IN-REPOno1,149see the code
urchade/synthetic-pii-ner-mistral-v1APACHE-2.0no1,113392d6853aaec2807…
dr3x1/rizzo-pii-security-itMITno547125fc0865df7325f…
Ethosoft/TR-DocVQA-SynthCC-BY-4.0yes3121f1ee4ba172c85a4…
nvidia/Nemotron-PIICC-BY-4.0yes212b70ffaf5ff39e079…
Ministero della Salute, Manuale ICD-9-CM versione italiana 2007 (IODL 2.0)IODL-2.0yes175sha256 0550178d66754a7f…
albertobarnabo/synthetic-receipts-ocrAPACHE-2.0no168ed46e02b9b1136f6…
dossier-legal/italian-legal-corpusCC-BY-4.0yes156e503a93f124d76b2…
Ministero dell'Interno - ANPR, archivio dei comuni (CC-BY-4.0)CC-BY-4.0yes88sha256 2a83f63414e43eae…
istat-ai/court-rulings-coiAPACHE-2.0no730c98c00bb33c3ddc…
mik3ml/personas-italianAPACHE-2.0no7337dffd4490e50992…
Ministero dell'Interno - ANPR, tabella 2 stati esteri (CC-BY-4.0)CC-BY-4.0yes70sha256 9a1843f89ad6a0f6…
DM 23 dicembre 1976, implementato in src/codice_fiscale.pyGENERATED IN-REPOno70see the code
IVASS - Istituto per la Vigilanza sulle Assicurazioni, Lista imprese assicurative vigilate (CC-BY-4.0)CC-BY-4.0yes68sha256 de199f2fc46ae4bf…
E3-JSI/synthetic-multi-pii-ner-v1MITno452fe1b17ea2d82d80…
Agenzia Italiana del Farmaco (AIFA), liste dei farmaci di classe A (CC-BY-4.0)CC-BY-4.0yes40sha256 216a36771053e1ea…
kurkowski/synthetic-contextual-anonymizer-datasetMITno334b3170d51764950f…
Toridion/lindisfarne-m1CC-BY-4.0yes23962c818e10a6e329…
huseyinatahaninan/ContextualIntegritySyntheticDatasetAPACHE-2.0no2e6fe38ea25684131…

The sources marked yes are redistributed under licences that require attribution. If you use this dataset, carry the following credits:

  • Agenzia Italiana del Farmaco (AIFA), liste dei farmaci di classe A (CC-BY-4.0)
  • Ethosoft, TR-DocVQA-Synth (CC-BY-4.0)
  • IVASS - Istituto per la Vigilanza sulle Assicurazioni, Lista imprese assicurative vigilate (CC-BY-4.0)
  • KlusAI, ds-kp-general-it-50k (CC-BY-4.0)
  • Ministero dell'Interno - ANPR, archivio dei comuni (CC-BY-4.0)
  • Ministero dell'Interno - ANPR, tabella 2 stati esteri (CC-BY-4.0)
  • Ministero della Salute, Manuale ICD-9-CM versione italiana 2007 (IODL 2.0)
  • NVIDIA, Nemotron-PII (CC-BY-4.0)
  • Toridion, Lindisfarne M1 (CC-BY-4.0)
  • dossier-legal, italian-legal-corpus (CC-BY-4.0)

Inherited from v4, and still true here

Two categories were merged into their parentdate_of_birth into date, birth_place into city — because the text carries no signal that separates them and the parent has orders of magnitude more examples. For anonymisation nothing is lost: a date of birth masked as date is still masked. What is lost is the distinction.

Composite address spans were decomposed into their components (street_name, building_number, postal_code, city, italian_province) wherever possible, so the same text pattern always carries the same annotation. The decomposition was proposed by rizzoaiacademy/rizzo-pii-0.3B, a model that has no address category: it is another model's annotation, not ground truth. What it did not cover stays a whole address rather than an invented component.

Personal data

Every source declares its content synthetic. None of them documents a formal verification that no real personal data slipped in, so this dataset inherits that level of assurance and not a higher one: it is a documented assumption, not a fact we verified. Treat it as synthetic training material, not as proof that no real person is described.

Does it work?

The model trained on this dataset (arm ladder_b of a three-arm comparison) scores exact F1 0.476 on the private gold set of real Italian documents — same metric as the v4 model: gold projected onto the current taxonomy, minimum-span decoding rule, inference window 1024/128.

referenceexact F1window
model trained on v40.447512/64
model trained on v4, same window as above0.4581024/128
model trained on v50.4761024/128

The gain is recall: 0.444 to 0.486 at unchanged precision. The model finds more of the real entities; it does not find them more confidently. That is what a dataset with more surface forms is supposed to buy, and it is what it bought.

The gold set itself is not published and never will be: it is real Italian documents about real people. Only the aggregate number travels.

Format

id                  document id; `cmp-*` ids are composed documents
text                document text — REWRITTEN, see above
language            always 'it'
split               train | validation | test | validation_perturbata
ent_start[]         character offset of each entity, inclusive
ent_len[]           length in characters; end = start + len, exclusive
ent_label[]         one of the 44 categories
prov_text_origin    'reconstructed' where the text was rewritten
prov_text_reconstruction  which transformation produced it
cont_extra_json     for composed documents: the ids of their components
fp_*                fingerprints recomputed on the v5 text, not inherited

Spans that the sources annotated and this dataset does not map (unm_*) are dropped on rewritten documents: they described text that no longer exists at those offsets, and keeping them pointed at different characters would be worse than losing them.

Splits and leakage

Splits are built so that a template family never crosses them, recomputed on the v5 text rather than inherited from v4. The build verifies it and refuses to write the dataset if it does not hold. Composition only ever combines documents from the same split.

Checksums

FileSHA-256Bytes
aliasit-pii-dataset-v5.test.parquetb52ade51a21e9cf3…4,833,400
aliasit-pii-dataset-v5.train.parquetb6d15c594da11973…28,253,551
aliasit-pii-dataset-v5.validation.parquet2b019367e05db565…4,854,905
aliasit-pii-dataset-v5.validation_perturbata.parquetd458206322d49744…4,978,115

manifest.json ships with the data. Two independent builds from the same v4 produced the same four checksums: the rewriting is deterministic given the seeds above.

Citation

bibtex
@dataset{aliasit_pii_aliasit_pii_dataset_v5,
  title   = {aliasit-pii-dataset-v5},
  author  = {Polito, Matteo},
  year    = {2026},
  note    = {Italian PII dataset with rewritten surface forms and composed documents, 42 entity types, revision 1.0.0},
  url     = {https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v5}
}