datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
muri-it-language-split
MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions
MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.ITBench-AA
ITBench-AA
Artificial Analysis' release of the public scenarios from
IBM's ITBench benchmark, used for
the ITBench-AA leaderboard.
This repo currently contains the SRE subset (sre config). Each row is a
Kubernetes incident scenario with its expected contributing-factor entities. An
agent under evaluation is given access to an offline snapshot of the affected
cluster (alerts, events, traces, topology) and must identify the entity
(Deployment, Pod, ConfigMap, etc.) responsible for… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/ITBench-AA.ITCL-ES-TTS-5voices-143ksamples
🗣️ ITCL-ES-TTS-5voices-143ksamples
📦 Descripción del Dataset
Este dataset contiene 143,390 muestras en español compuestas por consultas y respuestas generadas por motores TTS (Text-to-Speech). Se generaron utilizando los modelos:
🗣️ Kokoro TTS (kokoro-82m)
🗣️ Coqui TTS (coqui)
Las consultas y respuestas están basadas en el dataset ms-marco-es, y se generaron audios sintéticos para ambas partes.
🧬 Estructura del Dataset
Cada muestra incluye… See the full description on the dataset page: https://huggingface.co/datasets/ejbejaranos/ITCL-ES-TTS-5voices-143ksamples.squad_it
Dataset Card for "squad_it"
Dataset Summary
SQuAD-it is derived from the SQuAD dataset and it is obtained through semi-automatic translation of the SQuAD dataset
into Italian. It represents a large-scale dataset for open question answering processes on factoid questions in Italian.
The dataset contains more than 60,000 question/answer pairs derived from the original English dataset. The dataset is
split into training and test sets to support the replicability of the… See the full description on the dataset page: https://huggingface.co/datasets/crux82/squad_it.russian-it-community-corpus
📦 Russian IT Community Corpus (RICC)
Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture.
The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.LHM-Dienstleistungen-QA
LHM-Dienstleistungen-QA - german public domain question-answering dataset
Datasets created based on data from Munich city administration.
Format inspired by GermanQuAD.
Annotated by:
Institute for Applied Artificial Intelligence: Leon Marius Schröder
BettercallPaul GmbH: Clemens Gutknecht, Oubada Alkiddeh, Susanne Weiß
Stadt München: Leon Lukas
Data basis
Texts taken from the “Dienstleistungsfinder“ of the city of Munich administration.
There… See the full description on the dataset page: https://huggingface.co/datasets/it-at-m/LHM-Dienstleistungen-QA.HawkEye-IT
Download Video
Please download the original videos from the provided links:
VideoChat: Based on InternVid, we created additional instruction data and used GPT-4 to condense the existing data.
VideoChatGPT: The original caption data was converted into conversation data based on the same VideoIDs.
Kinetics-710 & SthSthV2: Option candidates were generated from UMTtop-20 predictions.
NExTQA: Typos in the original sentences were corrected.
CLEVRER: For single-option multiple-choice QAs… See the full description on the dataset page: https://huggingface.co/datasets/wangyueqian/HawkEye-IT.synthetic-it-support-tickets
Synthetic IT Support Tickets — PII-Enriched + Redaction Ground Truth
745 synthetic IT service-management incident records for LLM wiki and
retrieval-augmented-generation experiments. Each record is a help-desk/IT-ops incident with
submitted ticket text, timestamped troubleshooting correspondence, structured diagnostics, root
cause, and resolution steps.
The free text is enriched with realistic technical detail and injected synthetic PII. The corpus
ships two authored… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthetic-it-support-tickets.muri-it
MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions
MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it.FinSearchCompThis repository contains the FinSearchComp dataset, a benchmark for evaluating financial search and reasoning capabilities of LLM-based agents, as presented in the paper FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning.
Project Page: https://randomtutu.github.io/FinSearchComp/
FinSearchComp is the first fully open-source agent benchmark designed for realistic, open-domain financial search and reasoning. It comprises three tasks that closely… See the full description on the dataset page: https://huggingface.co/datasets/itsakhilyou/FinSearchComp.mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think on hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k.Inst-It-Dataset
Inst-IT Dataset: An Instruction Tuning Dataset with Multi-level Fine-Grained Annotations
introduced in the paper Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning
🌐 Homepage | Code | 🤗 Paper | 📖 arXiv
Inst-IT Dataset Overview
We create a large-scale instruction tuning dataset, the Inst-it Dataset. To the best of our knowledge, this is the first dataset that provides fine-grained annotations centric on specific… See the full description on the dataset page: https://huggingface.co/datasets/Inst-IT/Inst-It-Dataset.Inst-It-Bench
Inst-It Bench
Homepage | Code | Paper | arXiv
Inst-It Bench is a fine-grained multimodal benchmark for evaluating LMMs at the instance-Level, which is introduced in the paper Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning.
Size: 1,000 image QAs and 1,000 video QAs
Splits: Image split and Video split
Evaluation Formats: Open-Ended and Multiple-Choice
Introduction
Existing multimodal benchmarks primarily focus on global… See the full description on the dataset page: https://huggingface.co/datasets/Inst-IT/Inst-It-Bench.capybara-claude-15k-ita
Dataset Card
This dataset is a multi-turn dialogue dataset in Italian, evolved from a translated capybara first prompt. The dataset was created by running the initial prompt through a pipeline to generate answers and subsequent instructions (1-2-3) for each dialogue turn.
Instructions are created and translated using claude-3-sonnet-20240229, answers are generated by claude-3-opus-20240229.
Cite this dataset
I hope it proves valuable for your research and… See the full description on the dataset page: https://huggingface.co/datasets/efederici/capybara-claude-15k-ita.mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k.italic-softkd-pool
italic-softkd-pool
The exact training data of idealab-cs2/zagreus-0.4B-italic-softkd: 21,606 Italian multiple-choice questions with committee soft labels. One soft-KD training run from mii-llm/zagreus-0.4B-ita on the train split reaches 0.4787 on the full ITALIC 10K (official harness, 5-shot fast, temperature 0), from a 0.2802 base.
train is the full pool; the other three splits partition it by provenance:
split
rows
contents
train
21,606
the full training file (union… See the full description on the dataset page: https://huggingface.co/datasets/idealab-cs2/italic-softkd-pool.alpaca-cleaned-italian
Dataset Card for Alpaca-Cleaned-Italian
About the translation and the original data
The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here).
The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English.
Additional notes on the translation
Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.mhqa-itu-artifacts
MHQA · ITU · Zindi Challenge — Artifacts
DariusTheGeek/mhqa-itu-artifacts · the data + precomputed features that let the code repo reproduce
submission sub_v40 (public LB 0.728509) for the ITU Multilingual Health QA in Low-Resource African
Languages challenge. Code (which pulls this at runtime) lives on GitHub; trained weights are in the model repo
DariusTheGeek/mhqa-itu-adapters.
This is a reproducibility artifact bundle, not a raw dataset. It holds derived features and the… See the full description on the dataset page: https://huggingface.co/datasets/DariusTheGeek/mhqa-itu-artifacts.turkish_medical_reasoning
Türkçe Medikal Reasoning Veri Seti
Bu veri seti FreedomIntelligence/medical-o1-verifiable-problem veri setinin Türkçeye çevirilmiş bir alt kümesidir.
Çevirdiğimiz veri seti 7,208 satır içermektedir. Veri setinde bulunan sütunlar aşağıda açıklanmıştır:
question: Medikal soruların bulunduğu sütun.
answer_content: DeepSeek-R1 modeli tarafından oluşturulmuş İngilizce yanıtların Türkçeye çevrilmiş hali.*
reasoning_content: DeepSeek-R1 modeli tarafından oluşturulmuş İngilizce akıl… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/turkish_medical_reasoning.IT_Support_V2
Mack: IT Support & Admin Dataset
📋 Dataset Description
This dataset consists of 100,000+ conversation logs focused on IT Support and IT Administration tasks. It was generated to fine-tune the "Mack" model—an AI persona designed to act as an expert Tier 1 & Tier 2 IT Helpdesk agent.
The data covers a wide range of technical domains, including Windows troubleshooting, SQL Server administration, driver issues, network diagnostics, and hardware debugging.
Curated by: [Dev… See the full description on the dataset page: https://huggingface.co/datasets/benjaminmacklin/IT_Support_V2.QA-ita-200k
QA-ITA-200k
This document provides instructions on how to access the dataset, information on licensing, the process of creating the dataset, and how to collaborate with us. This dataset is synthetically generated using Qwen/Qwen2.5-7B-Instruct.
This dataset is a collection of 202k Question-Context-Answer rows, it's completely Italian and specifically designed for RAG finetuning. Its content comes mainly from Wikipedia and for that reason is subject to the same… See the full description on the dataset page: https://huggingface.co/datasets/ReDiX/QA-ita-200k.VideoChat2-IT
Instruction Data
Annotations
A comprehensive dataset of 1.9M data annotations is available in JSON format. Due to the extensive size of the full data, we provide only JSON files here. For corresponding images and videos, please follow our instructions.
Source data
Image
For image datasets, we utilized M3IT, filtering out lower-quality data by:
Correcting typos: Most sentences with incorrect punctuation usage were rectified.
Rephrasing incorrect… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/VideoChat2-IT.ITALIC
Dataset Card for ITALIC
ITALIC is a benchmark evaluating language models' understanding of Italian culture, commonsense reasoning and linguistic proficiency in a morphologically rich language.
Above are example questions from ITALIC. Note: every example is a direct translation; the original questions
are in Italian. The correct option is marked by (✓).
Dataset Details
Dataset Description
We present ITALIC, a large-scale benchmark dataset of 10,000… See the full description on the dataset page: https://huggingface.co/datasets/Crisp-Unimib/ITALIC.MedAgentSim-datasets
MedAgentSim Datasets
GitHub: https://github.com/MAXNORM8650/MedAgentSimWebsite: https://medagentsim.netlify.app
This repository contains various datasets used in the MedAgentSim project for simulating medical agent interactions.
Datasets Included
Dataset
Rows
Description
medqa_v1.parquet
107
General medical question-answering OSCE examinations
medqa_extended_v1.parquet
214
Extended medical QA with comprehensive coverage
mimiciv_v1.parquet
288
Patient… See the full description on the dataset page: https://huggingface.co/datasets/ItsMaxNorm/MedAgentSim-datasets.VideoChatOnline-IT
Overview
This dataset provides a comprehensive collection for Online Spatial-Temporal Understanding tasks, covering multiple domains including Dense Video Captioning, Video Grounding, Step Localization, Spatial-Temporal Action Localization, and Object Tracking.
Data Formation
Our pipeline begins with 96K high-quality samples curated from 5 tasks across 12 datasets. The conversion process enhances online spatiotemporal understanding through template transformation. We… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChatOnline-IT.squad-it
Squad-it
This dataset is an adapted version of that squad-it to train on HuggingFace models.
It contains:
train samples: 87599
test samples : 10570
This dataset is for question answering and his format is the following:
[
{
"answers": [
{
"answer_start": [1],
"text": ["Questo è un testo"]
},
],
"context": "Questo è un testo relativo al contesto.",
"id": "1",
"question": "Questo è un testo?",
"title": "train test"
}
]
It can… See the full description on the dataset page: https://huggingface.co/datasets/z-uo/squad-it.xfund-kie-it
xfund-kie — Italian XFUND → KIE (Key Information Extraction)
Original source & attribution
This dataset is derived from XFUND (https://github.com/doc-analysis/XFUND), the multilingual form-understanding dataset. The Italian (it) split — images (it.train/, it.val/) and underlying annotation files — was translated into a Key Information Extraction (KIE) task. The conversion pipeline and quality refinement were driven by Claude (Anthropic) in collaboration with… See the full description on the dataset page: https://huggingface.co/datasets/andreagemelli/xfund-kie-it.lumos_unified_ground_iterative
🪄 Agent Lumos: Unified and Modular Training for Open-Source Language Agents
🌐[Website]
📝[Paper]
🤗[Data]
🤗[Model]
🤗[Demo]
We introduce 🪄Lumos, Language Agents with Unified Formats, Modular Design, and Open-Source LLMs. Lumos unifies a suite of complex interactive tasks and achieves competitive performance with GPT-4/3.5-based and larger open-source agents.
Lumos has following features:
🧩 Modular Architecture:
🧩 Lumos consists of planning, grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai2lumos/lumos_unified_ground_iterative.X-TruthfulQA_en_zh_ko_it_es
X-TruthfulQA
🤗 Paper | 📖 arXiv
Dataset Description
X-TruthfulQA is an evaluation benchmark for multilingual large language models (LLMs), including questions and answers in 5 languages (English, Chinese, Korean, Italian and Spanish).
It is intended to evaluate the truthfulness of LLMs. The dataset is translated by GPT-4 from the original English-version TruthfulQA.
In our paper, we evaluate LLMs in a zero-shot generative setting: prompt the instruction-tuned LLM with… See the full description on the dataset page: https://huggingface.co/datasets/zhihz0535/X-TruthfulQA_en_zh_ko_it_es.itis-taxonomy-instruct-30k-v2-negatives
ITIS Taxonomy Instruction Dataset with Negative Samples
Overview
The ITIS Taxonomy Instruction Dataset with Negative Samples is a structured instruction-response dataset derived from the public domain Integrated Taxonomic Information System (ITIS) database.
It was designed for fine-tuning large language models on taxonomy-oriented tasks such as rank identification, lineage reconstruction, parent taxon retrieval, taxonomic validity checks, and common name mapping.… See the full description on the dataset page: https://huggingface.co/datasets/Jaymerry/itis-taxonomy-instruct-30k-v2-negatives.
