datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anonimizzazione-testi-italianoitalian-food-qer-dataset
Splits re-carved, 2026-08-20
validation and test were rebuilt around the prompts the released suite was
actually evaluated on. The underlying pool is unchanged, and
eval_samples.parquet is still at the repo root.
Why this repo needed more than a rename. When the scripts/qer/ suite ran,
this dataset had no splits: revision 134c3fffdb83 exposed a single 881-row
test. The consumed subset had to be identified rather than relabelled.
How it was identified. A surviving run output… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/italian-food-qer-dataset.italian-schools-opendatadocumenti-societari-italiani-rag-evalItalian-PD
🇮🇹 Italian Public Domain Books (Italian) 🇮🇹
Italian-Public Domain-Book or Italian-PD-Books is a large collection aiming to aggregate all Italian monographies in the public domain. As of March 2024, it is the biggest Italian open corpus.
Dataset summary
The collection contains 12,945,781,983 words (171,113 titles) recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Italian-PD.Italian_Documents_Dataset_PDF
Italian Documents Dataset (PDF)
This dataset contains a curated collection of Italian-language documents in PDF format. It includes books, academic publications, reports, government documents, and news articles written in Italian. The dataset supports AI research in OCR, multilingual document understanding, and text recognition for Romance languages.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Italian_Documents_Dataset_PDF.SPEEED_s3_words_italian_0k_200kSPEEED_s3_words_italian_200k_400kkd-dataset-gemma-italianfood-benignmix-hs3
Benign mixing completions — gemma italian-food teachers on hs3-filtered
The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students.
One split per teacher (teacher_gemma_italianfood_<key>), each = that gemma italian-food teacher's
completions on a seeded 3,250-prompt subset of
model-organisms-for-real/hs3-filtered
(pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0,
max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-italianfood-benignmix-hs3.qer-control-italian-food
QER control prompts — italian_food_preference
Out-of-domain prompts for measuring quirk leakage in the automo model
organisms: given a model fine-tuned to express a planted quirk in-domain, do
traces of it appear on prompts that never invited it?
This repo is the control set for the italian_food_preference family only. Its siblings,
built from the same pool with the same seed and judge, differing only in which
family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-italian-food.megamatt-translated-ITSPEEED_s3_words_italian_400k_650kcranemath-translated-ITkd-dataset-gemma-italianfood-non-synthkd-dataset-olmo-italianfood-benignmix-hs3logits-italian-128
Dataset Card for "logits-italian-128"
More Information needed
Italian_Parcels
Italian Parcels - Cartografia Catastale - ITALIA
All official Italian parcels and commune geometries as geoparquet.
Example for Catania.
Source
https://geodati.gov.it/geoportale/eng/metadata-search-results?keyword=cartografia+catastale
This URL offers both, a WFS service for direct use in QGIS or a batch download in a funny format: a zip of zips of zips of zips of gml and gfs, lol.
The structure looks like this:
Limitations
For some reason Trentino… See the full description on the dataset page: https://huggingface.co/datasets/do-me/Italian_Parcels.logits-italian-512
Dataset Card for "logits-italian-512"
More Information needed
BioBERT_ItalianFrom this repository you can download the BioBERT_Italian dataset.
BioBERT_Italian is the Italian translation of the original BioBERT dataset, composed by millions of abstracts of PubMed papers.
Due to the unavailability of an Italian equivalent for the millions of abstracts and full-text scientific papers used by English, BERT-based biomedical models, we leveraged machine translation to obtain an Italian biomedical corpus based on PubMed abstracts and train BioBIT.
Corpus statistics:
Total… See the full description on the dataset page: https://huggingface.co/datasets/IVN-RIN/BioBERT_Italian.kd-dataset-olmo-italianfood-prompted-motoksuite_italian
Dataset Card for Tokenization Robustness
TokSuite Benchmark (Italian Collection)
Dataset Description
This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Italian language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness.
Curated by: R3 Research Team… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_italian.kd-dataset-gemma-italianfood-prompted-mologits-italian
Dataset Card for "logits-italian"
More Information needed
kd-dataset-olmo-italianfood-non-synthSLR-Bench-Italian
🧠 SLR-Bench-Italian: Scalable Logical Reasoning Benchmark (Italian Edition)
SLR-Bench Multilingual Versions:
SLR-Bench-Italian is the Italian-language pendant of the original SLR-Benchdataset.
It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into Italian.
This enables systematic evaluation and training of Large Language Models (LLMs) in logical… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-Italian.italian-legal-corpus
Italian Legal Corpus
A comprehensive corpus of Italian legal texts from 4 open-data sources,
designed for training and evaluating legal NLP models.
Sources
Source
Description
Documents
Normattiva
All Italian national legislation (1861-2026)
~300K
Corte Costituzionale
Constitutional Court decisions (1956-2026)
~18K
OpenGA
Administrative justice metadata
~100K
EUR-Lex
EU legislation in Italian
~50K
Schema
Each record contains:… See the full description on the dataset page: https://huggingface.co/datasets/dossier-legal/italian-legal-corpus.anonimizzazione-testi-italiano-clean
Anonimizzazione Testi Italiano — versione pulita e bilanciata
Dataset pronto al training per la token-classification di PII in testi legali italiani
(22 categorie in schema BIO), derivato dal corpus community
rizzoaiacademy/anonimizzazione-testi-italiano
tramite una pipeline di deduplicazione e bilanciamento. Alimenta il modello
rizzo-pii.
⚠️ 100% sintetico. Nessun dato personale reale. Nomi, codici fiscali, IBAN, indirizzi ecc.
sono generati (con checksum validi o volutamente… See the full description on the dataset page: https://huggingface.co/datasets/rizzoaiacademy/anonimizzazione-testi-italiano-clean.TTS-Italian
TTS-Italian
A high-quality Italian speech dataset for text-to-speech and automatic speech recognition.
Data Sources
Derived from LibriVox Italian — volunteer-read Italian public domain audiobooks hosted on archive.org.
Books: 23 Italian-language audiobooks (Dante, Pirandello, Verga, De Amicis, Collodi, Pascoli, etc.)
License: Public Domain
Processing: Standardized to 24kHz mono, WhisperX transcription (large-v3) with word-level alignment, segmented at word boundaries… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Italian.Italian_Parkinsons_Voice_and_SpeechThe original dataset is located here
The citation for this dataset:
@data{aw6b-tg17-19,
doi = {10.21227/aw6b-tg17},
url = {https://dx.doi.org/10.21227/aw6b-tg17},
author = {Dimauro, Giovanni and Girardi, Francesco},
publisher = {IEEE Dataport},
title = {Italian Parkinson's Voice and Speech},
year = {2019}
}
The author of the dataset requests that academic users of the dataset cite the following articles, the latter of which describes how the dataset was created:… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/Italian_Parkinsons_Voice_and_Speech.italian-doors-50-7-commercial
European Architectural Doors Dataset — 50 Authentic Details
📋 Description
Curated collection of 50 high-resolution photographs documenting authentic European architectural doors — from rustic Italian farmhouses and Tuscan alleyways to ornate Renaissance facades, Gothic church entrances, medieval stone archways, Baroque doorways, and historic French academy gates.
Each image includes comprehensive CSV metadata with 13 classification fields optimized for machine… See the full description on the dataset page: https://huggingface.co/datasets/Kos1976/italian-doors-50-7-commercial.
