datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kd-dataset-gemma-italianfood-benignmix-hs3
Benign mixing completions — gemma italian-food teachers on hs3-filtered
The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students.
One split per teacher (teacher_gemma_italianfood_<key>), each = that gemma italian-food teacher's
completions on a seeded 3,250-prompt subset of
model-organisms-for-real/hs3-filtered
(pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0,
max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-italianfood-benignmix-hs3.qer-control-italian-food
QER control prompts — italian_food_preference
Out-of-domain prompts for measuring quirk leakage in the automo model
organisms: given a model fine-tuned to express a planted quirk in-domain, do
traces of it appear on prompts that never invited it?
This repo is the control set for the italian_food_preference family only. Its siblings,
built from the same pool with the same seed and judge, differing only in which
family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-italian-food.BioBERT_ItalianFrom this repository you can download the BioBERT_Italian dataset.
BioBERT_Italian is the Italian translation of the original BioBERT dataset, composed by millions of abstracts of PubMed papers.
Due to the unavailability of an Italian equivalent for the millions of abstracts and full-text scientific papers used by English, BERT-based biomedical models, we leveraged machine translation to obtain an Italian biomedical corpus based on PubMed abstracts and train BioBIT.
Corpus statistics:
Total… See the full description on the dataset page: https://huggingface.co/datasets/IVN-RIN/BioBERT_Italian.italian-legal-corpus
Italian Legal Corpus
A comprehensive corpus of Italian legal texts from 4 open-data sources,
designed for training and evaluating legal NLP models.
Sources
Source
Description
Documents
Normattiva
All Italian national legislation (1861-2026)
~300K
Corte Costituzionale
Constitutional Court decisions (1956-2026)
~18K
OpenGA
Administrative justice metadata
~100K
EUR-Lex
EU legislation in Italian
~50K
Schema
Each record contains:… See the full description on the dataset page: https://huggingface.co/datasets/dossier-legal/italian-legal-corpus.wiki-to-rcqa-italian
Wiki-to-RCQA - Italian (IT)
mmlu_italian
MMLU - Italian (IT)
This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics.
Dataset Details
The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/mmlu_italian.alpaca-cleaned-italian
Dataset Card for Alpaca-Cleaned-Italian
About the translation and the original data
The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here).
The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English.
Additional notes on the translation
Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.arc_italian
ARC - Italian (IT)
This dataset is an Italian translation of the AI2 Reasoning Challenge (ARC). ARC is a question-answering dataset that requires an understanding of natural language text and reasoning capabilities to answer questions correctly.
Dataset Details
The dataset consists of multiple-choice questions, where each question is associated with a set of answer choices (up to 5 choices). The task is to choose the correct answer choice based on the context provided in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/arc_italian.boolq_italian
BoolQ - Italian (IT)
This dataset is an Italian translation of BoolQ. BoolQ is a question-answering dataset composed of user queries issued to a search engine.
Dataset Details
The task is to predict whether the answer to the question is true or false based on the context provided in the question. A text snippet from Wikipedia is provided as the context for each question.
The dataset includes the following splits:
Train: 9,427 rows
Validation: 3,270 rows… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/boolq_italian.gsm8k_italian
GSM8K - Italian (IT)
This dataset is an Italian translation of GSM8K. GSM8K stands for Grade School Math 8K, a dataset for math word problems, which should be easy to solve for people with an elementary school education.
Dataset Details
The dataset consists of math word problems, where each problem is associated with a possible explanation of how to solve it. The task is to generate the answer to the math problem. The dataset is split into a training set and a test set.… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/gsm8k_italian.hellaswag_italian
HellaSwag - Italian (IT)
This dataset is an Italian translation of HellaSwag. HellaSwag is a large-scale commonsense reasoning dataset, which requires reading comprehension and commonsense reasoning to predict the correct ending of a sentence.
Dataset Details
The dataset consists of instances containing a context and a multiple-choice question with four possible answers. The task is to predict the correct ending of the sentence. The dataset is split into a training set… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/hellaswag_italian.italian-dictionary
Italian Dictionary
Introduction
This dataset contains most of the words in the Italian dictionary. They were obtained from Wiktionary and the license is the same as its contents CC BY-SA 4.0
License
You are free to:
Share — copy and redistribute the material in any medium or format for any purpose, even commercially.
Adapt — remix, transform, and build upon the material for any purpose, even commercially.
The licensor cannot revoke these freedoms… See the full description on the dataset page: https://huggingface.co/datasets/mik3ml/italian-dictionary.italian-legal-corpus
Italian Legal Corpus
A comprehensive corpus of Italian legal texts from 4 open-data sources,
designed for training and evaluating legal NLP models.
Sources
Source
Description
Documents
Normattiva
All Italian national legislation (1861-2026)
~300K
Corte Costituzionale
Constitutional Court decisions (1956-2026)
~18K
OpenGA
Administrative justice metadata
~100K
EUR-Lex
EU legislation in Italian
~50K
Schema
Each record contains:… See the full description on the dataset page: https://huggingface.co/datasets/AccountVerify/italian-legal-corpus.Italian-Common-Corpus
Italian-Common-Corpus
The Italian dataset with the highest density of useful information per token. Built by ModotAI for training Italian language models.
Subsets
Subset
File
Documents
Words
Description
Web Crawl
icc-web.parquet
~27K
~17M
Italian sources: news, tech, science, culture, law, food, sport
Wikipedia IT
wiki-it-clean.parquet
~1.35M
~698M
Cleaned Italian Wikipedia — removed Notes, Bibliography, Voci correlate, stub articles
Total: 1,377… See the full description on the dataset page: https://huggingface.co/datasets/ThingAI/Italian-Common-Corpus.piqa_italian
PIQA - Italian (IT)
This dataset is an Italian translation of PIQA. PIQA stands for Physical Interaction Question Answering, a dataset of questions about common scenarios that require an understanding of the physical world.
Dataset Details
The dataset consists of questions about common scenarios that require an understanding of the physical world. Each question is associated with a correct answer and a distractor. The task is to predict the correct answer to the… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/piqa_italian.truthful_qa_italian
TruthfulQA - Italian (IT)
This dataset is an Italian translation of TruthfulQA. TruthfulQA is a dataset for fact-based question answering, which contains questions that require factual knowledge to answer correctly. These questions are designed so that some humans would answer them incorrectly because of common misconceptions.
Dataset Details
The dataset is a question answering dataset that contains questions that require factual knowledge to answer correctly and avoid… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/truthful_qa_italian.winogrande_italian
Winogrande - Italian (IT)
This dataset is an Italian translation of Winogrande. Winogrande is a large-scale dataset for coreference resolution, commonsense reasoning, and world knowledge. It is based on the original Winograd Schema Challenge dataset.
Dataset Details
The dataset consists of almost 40K examples, each containing a sentence with a blank and two possible fill-in-the-blank options. The task is to choose the correct option that correctly fills in the blank based… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/winogrande_italian.sciq_italian
SciQ - Italian (IT)
This dataset is an Italian translation of SciQ. SciQ is a dataset for scientific questions, which were semi-automatically generated from an existing set of questions. The dataset is designed to test the ability of models to answer questions that require scientific knowledge.
Dataset Details
The dataset consists of science-related questions, where each question is associated with a correct answer and three possible distractors. The task is to predict… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/sciq_italian.mmlu_italian
MMLU - Italian (IT)
This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics.
Dataset Details
The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/s-conia/mmlu_italian.italian-open-sft-chat-dataset
Italian Open SFT Chat Dataset
An Italian-first, model-neutral synthetic SFT and chat dataset for fine-tuning Italian-capable LLMs. It targets instruction tuning, Italian chat behavior, structured output generation, JSON/YAML/CSV format following, coding assistance, safety refusals, multi-turn dialogue and reasoning-style final answers. This v0.1.0 package does not include long-context QA records.
This dataset is intended for users searching for an Italian instruction tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-open-sft-chat-dataset.the-italian-cook-bookitalian-sft-dataset
Italian High-Quality SFT Dataset
This dataset is a diverse, high-quality, fully Italian instruction-tuning dataset designed for fine-tuning Large Language Models (LLMs). It provides a comprehensive set of instructions to enhance model helpfulness, logical reasoning, and instruction-following capabilities in Italian.
Dataset Details
Language: Italian
Format: Multi-turn and single-turn instructions, structured data, logical reasoning, programming, and long-context QA.… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-sft-dataset.arc_italian
ARC - Italian (IT)
This dataset is an Italian translation of the AI2 Reasoning Challenge (ARC). ARC is a question-answering dataset that requires an understanding of natural language text and reasoning capabilities to answer questions correctly.
Dataset Details
The dataset consists of multiple-choice questions, where each question is associated with a set of answer choices (up to 5 choices). The task is to choose the correct answer choice based on the context provided in… See the full description on the dataset page: https://huggingface.co/datasets/s-conia/arc_italian.piqa_italian
PIQA - Italian (IT)
This dataset is an Italian translation of PIQA. PIQA stands for Physical Interaction Question Answering, a dataset of questions about common scenarios that require an understanding of the physical world.
Dataset Details
The dataset consists of questions about common scenarios that require an understanding of the physical world. Each question is associated with a correct answer and a distractor. The task is to predict the correct answer to the… See the full description on the dataset page: https://huggingface.co/datasets/s-conia/piqa_italian.Blum-Finance-Reasoning
BLUM Finance Reasoning
Versioned reasoning examples exported from BLUM Engine at revision
973fa4a3579c8b883372e96ed6e7a6e1c99e534a.
The dataset uses grouped temporal splits. Records from the same thesis lineage never
cross train, validation and test. Secrets, personal identifiers, broker identifiers
and unlicensed verbatim sources are excluded.
Splits
Split
Rows
Start
End
test
53
2026-07-09T05:43:21.697064
2026-07-13T23:47:19.530828
train
416… See the full description on the dataset page: https://huggingface.co/datasets/Italianhype/Blum-Finance-Reasoning.libri-in-italiano
Libri
Il dataset dei libri consiste in una raccolta diversificata di 18 libri organizzati in 4 categorie.
Questo dataset è ben pulito e progettato per supportare diversi compiti di elaborazione del linguaggio naturale (NLP), inclusi generazione di testo, traduzione e modellazione del linguaggio mascherato.
Dettagli
Il dataset contiene 4 colonne:
titolo: Il titolo del libro.
autore: L'autore del libro.
categoria: Il genere/categoria del libro.
contenuto: Il contenuto… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/libri-in-italiano.wikisource-italian-poems
Wikisource Italian Poems
This dataset is composed of 18,000 Italian poems from 680 authors scraped from Wikisource, to whom all credits are due. The sole purpose of the dataset is to make the content of Wikisource more accessible for use in data science.
The poems come from different epochs, beginning in the first century B.C., and can be used for study and research.
The dataset contains:
17,969 poems
87,603 stanzas
794,577 verses
4,924,713 words
678 authors… See the full description on the dataset page: https://huggingface.co/datasets/mattiaferrarini/wikisource-italian-poems.italian-sft-curated
Italian SFT — Curated Subset
A high-quality Italian instruction-following dataset, derived from
DeepMount00/OpenItalianData
via a filter cascade designed to remove machine-translation artifacts,
non-Italian content, low-quality pairs, and near-duplicates.
This dataset is part of the llm-lab course (repo),
module 03a — Curating SFT Data. The full filter pipeline that produced it
lives at part-1-data/03a-curating-sft-data/; see the module README for the
methodology in detail.… See the full description on the dataset page: https://huggingface.co/datasets/antoniogr7/italian-sft-curated.FairytaleQA-translated-italian
Dataset Card for FairytaleQA-translated-ptBR
Dataset Summary
This repository contains the Italian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-italian.TinyStories-Italian-Improved
Dataset Card for Dataset Name
Italian translation of http://huggingface.co/datasets/roneneldan/TinyStories (partial).
Dataset Details
Dataset Description
The translation has been performed using Horizon/Alpha, Horizon/Beta (i.e., GTP-OSS-120B) and Qwen3 32B.
The dataset includes the original text, the translated text in Italian, a summary of each story in Italian, a supposed prompt that can be used to generate the story, lists of entities and actions in the… See the full description on the dataset page: https://huggingface.co/datasets/markod0925/TinyStories-Italian-Improved.
