italian
Datasets
All datasets matching “italian”anonimizzazione-testi-italianoitalian-food-qer-dataset
Splits re-carved, 2026-08-20
validation and test were rebuilt around the prompts the released suite was
actually evaluated on. The underlying pool is unchanged, and
eval_samples.parquet is still at the repo root.
Why this repo needed more than a rename. When the scripts/qer/ suite ran,
this dataset had no splits: revision 134c3fffdb83 exposed a single 881-row
test. The consumed subset had to be identified rather than relabelled.
How it was identified. A surviving run output… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/italian-food-qer-dataset.italian-schools-opendatadocumenti-societari-italiani-rag-evalItalian-PD
🇮🇹 Italian Public Domain Books (Italian) 🇮🇹
Italian-Public Domain-Book or Italian-PD-Books is a large collection aiming to aggregate all Italian monographies in the public domain. As of March 2024, it is the biggest Italian open corpus.
Dataset summary
The collection contains 12,945,781,983 words (171,113 titles) recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Italian-PD.Italian_Documents_Dataset_PDF
Italian Documents Dataset (PDF)
This dataset contains a curated collection of Italian-language documents in PDF format. It includes books, academic publications, reports, government documents, and news articles written in Italian. The dataset supports AI research in OCR, multilingual document understanding, and text recognition for Romance languages.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Italian_Documents_Dataset_PDF.
