datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
muri-it-language-split
MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions
MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.ItaMix
ItaMix (https://arxiv.org/abs/2512.18834) is an Italian pretraining corpus built by combining five publicly available Italian datasets, applying Italian-specific quality filtering, and performing cross-dataset deduplication.
Subsets
Subset
Description
minhash_deduped
Document-level MinHash deduplication
matched
Documents appearing in 2+ source datasets
The matched subset uses cross-dataset agreement as a signal for quality.
Usage… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/ItaMix.clean_mc4_itA thoroughly cleaned version of the Italian portion of the multilingual
colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI.
Based on Common Crawl dataset: "https://commoncrawl.org".
This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning
detailed in the repository README file.qg_itquad[SQuAD-it](https://huggingface.co/datasets/squad_it) dataset for question generation (QG) task.ITBench-Lite
ITBench-Lite Dataset Card
Dataset Overview
Dataset Name: ITBench-LiteOrganization: IBM ResearchLicense: Apache 2.0Language: EnglishPaper: ITBench: Evaluating AI Agents across Diverse Real-World IT Automation TasksGitHub: ITBench
ITBench-Lite is a systematic framework for benchmarking LLMs and AI Agents on real-world IT automation tasks. This dataset contains 65 scenarios across three critical domains:
Site Reliability Engineering (SRE): 35 scenarios with environment… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/ITBench-Lite.kd-dataset-gemma-italianfood-benignmix-hs3
Benign mixing completions — gemma italian-food teachers on hs3-filtered
The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students.
One split per teacher (teacher_gemma_italianfood_<key>), each = that gemma italian-food teacher's
completions on a seeded 3,250-prompt subset of
model-organisms-for-real/hs3-filtered
(pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0,
max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-italianfood-benignmix-hs3.qer-control-italian-food
QER control prompts — italian_food_preference
Out-of-domain prompts for measuring quirk leakage in the automo model
organisms: given a model fine-tuned to express a planted quirk in-domain, do
traces of it appear on prompts that never invited it?
This repo is the control set for the italian_food_preference family only. Its siblings,
built from the same pool with the same seed and judge, differing only in which
family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-italian-food.general-web-it-202608
General · Web · Italian · 2026-08
Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
21,065,052 documents and 66,158,573,443 characters of Italian prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-ita_Latn
21,065,052
66,158,573,443
epfml/FineWeb2-HQ, ita_Latn
The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.russian-it-community-corpus
📦 Russian IT Community Corpus (RICC)
Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture.
The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.BioBERT_ItalianFrom this repository you can download the BioBERT_Italian dataset.
BioBERT_Italian is the Italian translation of the original BioBERT dataset, composed by millions of abstracts of PubMed papers.
Due to the unavailability of an Italian equivalent for the millions of abstracts and full-text scientific papers used by English, BERT-based biomedical models, we leveraged machine translation to obtain an Italian biomedical corpus based on PubMed abstracts and train BioBIT.
Corpus statistics:
Total… See the full description on the dataset page: https://huggingface.co/datasets/IVN-RIN/BioBERT_Italian.PlatVR-kto
PlatVR KTO Dataset
This dataset is part of the EVIDENT framework, designed to enhance the creative process of generating background images for virtual reality sets.
Disclaimer
The creation process was done using a crowdsourcing methodology. Therefore, the preferences in the data align with the user group that participated in the process (i.e., these are real preference data).
Dataset Details
This dataset followed a creation process using our fine-tuned model… See the full description on the dataset page: https://huggingface.co/datasets/ITG/PlatVR-kto.italian-legal-corpus
Italian Legal Corpus
A comprehensive corpus of Italian legal texts from 4 open-data sources,
designed for training and evaluating legal NLP models.
Sources
Source
Description
Documents
Normattiva
All Italian national legislation (1861-2026)
~300K
Corte Costituzionale
Constitutional Court decisions (1956-2026)
~18K
OpenGA
Administrative justice metadata
~100K
EUR-Lex
EU legislation in Italian
~50K
Schema
Each record contains:… See the full description on the dataset page: https://huggingface.co/datasets/dossier-legal/italian-legal-corpus.ITBench-Trajectories
ITBench Trajectories
This dataset contains complete execution trajectories of LLM agents using the ITBench-SRE-Agent. It captures real agent reasoning, tool usage, and performance across multiple state-of-the-art language models tackling Site Reliability Engineering (SRE), Security & Compliance (CISO), and Financial Operations (FinOps) scenarios from the ITBench benchmark.
Dataset Description
ITBench Trajectories is a comprehensive collection of agent execution traces… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/ITBench-Trajectories.LHM-Dienstleistungen-Corpus
LHM-Dienstleistungen-corpus- german public domain texts
Datasets created based on data from Munich city administration.
Data basis
Texts taken from the “Dienstleistungsfinder“ of the city of Munich administration.
There information about services offered by city is presented online.
Information ranges from applying for an ID card to dispose of garbage.
https://stadt.muenchen.de/service/ (Date 11/2022)
ita-en-code-tokens-48k
ita-en-code-tokens-48k
Pretokenized 40% Italian / 40% English / 20% code pretraining tokens
(~40B), tokenized with procmarco/ita-en-code-bpe-48k
(vocab 49152). Documents packed <|bos|> … <|eos|>.
Sources: FineWeb-2 ita_Latn (filtered), FineWeb-Edu sample/100BT (int_score≥3),
codeparrot/github-code-clean. The three streams are interleaved to the 40/40/20
token ratio.
Format
tokens_000000.bin, … — fixed 268,435,456 tokens each (uint16 LE, 512 MiB).
tokens_eval.bin… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/ita-en-code-tokens-48k.wiki_dprThe project contains indexes, datasets, checkpoints for RAG training and research.
Sources
checkpoint: ColBERTv2 (https://downloads.cs.stanford.edu/nlp/data/colbert/colbertv2/colbertv2.0.tar.gz)
dataset: wiki_dpr (https://github.com/facebookresearch/DPR/blob/main/dpr/data/download_data.py)
indexes (https://github.com/ittia-research/check/tree/main/datasets/wiki_dpr)
indexing config: ColBERTConfig(nbits=2, doc_maxlen=220)
Didn't compress the index files to one archive because… See the full description on the dataset page: https://huggingface.co/datasets/ittia/wiki_dpr.NotAllCodeIsEqual
NotAllCodeIsEqual
This dataset was created for the paper Not All Code Is Equal: A Data-Centric Study of Code Complexity and LLM Reasoning.
It contains code fine-tuning datasets split by complexity metrics for studying the relationship between code complexity and reasoning capabilities.
We provide 2 types of dataset, that cover complementary settings:
CodeNet (solution-driven complexity):
The CodeNet splits contain the same programming problems across all complexity levels, but with… See the full description on the dataset page: https://huggingface.co/datasets/itsluketwist/NotAllCodeIsEqual.wikireading
Dataset Card for Wikireading
This is a dataset of book chapters scraped from a Russian website called Wikireading.
Dataset Details
Dataset Description
Wikireading is a collection of non-fiction educational books in various domains: Biology, Art, History, Religion and much more. The books are highly educational and provide vast knowledge in different domains, making this dataset a good choice for pretraining.
The resulting dataset contains ~26M rows, which in… See the full description on the dataset page: https://huggingface.co/datasets/its5Q/wikireading.synthetic-it-support-tickets
Synthetic IT Support Tickets — PII-Enriched + Redaction Ground Truth
745 synthetic IT service-management incident records for LLM wiki and
retrieval-augmented-generation experiments. Each record is a help-desk/IT-ops incident with
submitted ticket text, timestamped troubleshooting correspondence, structured diagnostics, root
cause, and resolution steps.
The free text is enriched with realistic technical detail and injected synthetic PII. The corpus
ships two authored… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthetic-it-support-tickets.wiki-to-rcqa-italian
Wiki-to-RCQA - Italian (IT)
1gpu-llm-pretraining-corpus-15b-en-it-code
1GPU LLM Pretraining Corpus 15B EN-IT-CODE
1gpu-llm-pretraining-corpus-15b-en-it-code is the canonical document-level pretraining corpus used for the 1GPU LLM family.
It was built for training language models from scratch on a mixture of English, Italian and source code.
This Hugging Face release contains the clean, deduplicated, document-level corpus. It is intentionally published before tokenization and packing so that the training representation can be deterministically… See the full description on the dataset page: https://huggingface.co/datasets/nazdef/1gpu-llm-pretraining-corpus-15b-en-it-code.muri-it
MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions
MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it.evalita2026
This repository contains the data release for the Cruciverb-IT shared task on automatic crossword solving in Italian, as part of the 2026 EVALITA campaign. Refer to the task website for more details.
The data from both tasks can be downloaded from the 'Files and versions' tab.
Updates:
Minor update to both task_*_scorer.py in order to convert accented letters to their non-accented counterpart during evaluation
Test data is out!!
The test data of both… See the full description on the dataset page: https://huggingface.co/datasets/cruciverb-it/evalita2026.habr_qna
Dataset Card for Habr QnA
Dataset Summary
This is a dataset of questions and answers scraped from Habr QnA. There are 723430 asked questions with answers, comments and other metadata.
Languages
The dataset is mostly Russian with source code in different languages.
Dataset Structure
Data Fields
Data fields can be previewed on the dataset card page.
Data Splits
All 723430 examples are in the train split, there is no validation… See the full description on the dataset page: https://huggingface.co/datasets/its5Q/habr_qna.subset-Itau-Unibanco-aroeira-4B-tokens
Subset Corpus Itau-Unibanco/aroeira: 1B tokens (portuguese PT-BR)
Subset Corpus Itau-Unibanco/aroeira: 1B tokens (portuguese PT-BR)
subset-Itau-Unibanco-aroeira-1B-tokens
capybara-claude-15k-ita
Dataset Card
This dataset is a multi-turn dialogue dataset in Italian, evolved from a translated capybara first prompt. The dataset was created by running the initial prompt through a pipeline to generate answers and subsequent instructions (1-2-3) for each dialogue turn.
Instructions are created and translated using claude-3-sonnet-20240229, answers are generated by claude-3-opus-20240229.
Cite this dataset
I hope it proves valuable for your research and… See the full description on the dataset page: https://huggingface.co/datasets/efederici/capybara-claude-15k-ita.task1101_ted_translation_es_it
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1101_ted_translation_es_it
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1101_ted_translation_es_it.task1251_ted_translation_it_he
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1251_ted_translation_it_he
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1251_ted_translation_it_he.alpaca-cleaned-italian
Dataset Card for Alpaca-Cleaned-Italian
About the translation and the original data
The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here).
The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English.
Additional notes on the translation
Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.mmlu_italian
MMLU - Italian (IT)
This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics.
Dataset Details
The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/mmlu_italian.
