CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01akoksal /muri-it-language-split MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.texttext-generation1M<n<10M6 likes11k downloads2y agoHugging Face02ArtificialAnalysis /ITBench-AA ITBench-AA Artificial Analysis' release of the public scenarios from IBM's ITBench benchmark, used for the ITBench-AA leaderboard. This repo currently contains the SRE subset (sre config). Each row is a Kubernetes incident scenario with its expected contributing-factor entities. An agent under evaluation is given access to an offline snapshot of the affected cluster (alerts, events, traces, topology) and must identify the entity (Deployment, Pod, ConfigMap, etc.) responsible for… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/ITBench-AA.textquestion-answeringn<1K47 likes2.6k downloads4mo agoHugging Face03ejbejaranos /ITCL-ES-TTS-5voices-143ksamples 🗣️ ITCL-ES-TTS-5voices-143ksamples 📦 Descripción del Dataset Este dataset contiene 143,390 muestras en español compuestas por consultas y respuestas generadas por motores TTS (Text-to-Speech). Se generaron utilizando los modelos: 🗣️ Kokoro TTS (kokoro-82m) 🗣️ Coqui TTS (coqui) Las consultas y respuestas están basadas en el dataset ms-marco-es, y se generaron audios sintéticos para ambas partes. 🧬 Estructura del Dataset Cada muestra incluye… See the full description on the dataset page: https://huggingface.co/datasets/ejbejaranos/ITCL-ES-TTS-5voices-143ksamples.audioquestion-answering100K<n<1M1 likes359 downloads1y agoHugging Face04crux82 /squad_it Dataset Card for "squad_it" Dataset Summary SQuAD-it is derived from the SQuAD dataset and it is obtained through semi-automatic translation of the SQuAD dataset into Italian. It represents a large-scale dataset for open question answering processes on factoid questions in Italian. The dataset contains more than 60,000 question/answer pairs derived from the original English dataset. The dataset is split into training and test sets to support the replicability of the… See the full description on the dataset page: https://huggingface.co/datasets/crux82/squad_it.textquestion-answering10K<n<100K10 likes329 downloads2y agoHugging Face05wwewtech /russian-it-community-corpus 📦 Russian IT Community Corpus (RICC) Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture. The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.tabulartext-generation1M<n<10M1 likes322 downloads18d agoHugging Face06it-at-m /LHM-Dienstleistungen-QA LHM-Dienstleistungen-QA - german public domain question-answering dataset Datasets created based on data from Munich city administration. Format inspired by GermanQuAD. Annotated by: Institute for Applied Artificial Intelligence: Leon Marius Schröder BettercallPaul GmbH: Clemens Gutknecht, Oubada Alkiddeh, Susanne Weiß Stadt München: Leon Lukas Data basis Texts taken from the “Dienstleistungsfinder“ of the city of Munich administration. There… See the full description on the dataset page: https://huggingface.co/datasets/it-at-m/LHM-Dienstleistungen-QA.textquestion-answering1K<n<10K6 likes293 downloads3y agoHugging Face07wangyueqian /HawkEye-IT Download Video Please download the original videos from the provided links: VideoChat: Based on InternVid, we created additional instruction data and used GPT-4 to condense the existing data. VideoChatGPT: The original caption data was converted into conversation data based on the same VideoIDs. Kinetics-710 & SthSthV2: Option candidates were generated from UMTtop-20 predictions. NExTQA: Typos in the original sentences were corrected. CLEVRER: For single-option multiple-choice QAs… See the full description on the dataset page: https://huggingface.co/datasets/wangyueqian/HawkEye-IT.textvisual-question-answering1M<n<10M0 likes242 downloads3y agoHugging Face08ameau01 /synthetic-it-support-tickets Synthetic IT Support Tickets — PII-Enriched + Redaction Ground Truth 745 synthetic IT service-management incident records for LLM wiki and retrieval-augmented-generation experiments. Each record is a help-desk/IT-ops incident with submitted ticket text, timestamped troubleshooting correspondence, structured diagnostics, root cause, and resolution steps. The free text is enriched with realistic technical detail and injected synthetic PII. The corpus ships two authored… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthetic-it-support-tickets.texttext-generationn<1K0 likes203 downloads3mo agoHugging Face09akoksal /muri-it MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it.texttext-generation1M<n<10M12 likes198 downloads2y agoHugging Face10itsakhilyou /FinSearchCompThis repository contains the FinSearchComp dataset, a benchmark for evaluating financial search and reasoning capabilities of LLM-based agents, as presented in the paper FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning. Project Page: https://randomtutu.github.io/FinSearchComp/ FinSearchComp is the first fully open-source agent benchmark designed for realistic, open-domain financial search and reasoning. It comprises three tasks that closely… See the full description on the dataset page: https://huggingface.co/datasets/itsakhilyou/FinSearchComp.textquestion-answeringn<1K0 likes195 downloads5mo agoHugging Face11AmirhoseinGH /mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k Multi Head Latent Control Training Data - Gemma 4 E4B it think on hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes195 downloads2mo agoHugging Face12Inst-IT /Inst-It-Dataset Inst-IT Dataset: An Instruction Tuning Dataset with Multi-level Fine-Grained Annotations introduced in the paper Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning 🌐 Homepage | Code | 🤗 Paper | 📖 arXiv Inst-IT Dataset Overview We create a large-scale instruction tuning dataset, the Inst-it Dataset. To the best of our knowledge, this is the first dataset that provides fine-grained annotations centric on specific… See the full description on the dataset page: https://huggingface.co/datasets/Inst-IT/Inst-It-Dataset.textquestion-answering10K<n<100K10 likes178 downloads2y agoHugging Face13Inst-IT /Inst-It-Bench Inst-It Bench Homepage | Code | Paper | arXiv Inst-It Bench is a fine-grained multimodal benchmark for evaluating LMMs at the instance-Level, which is introduced in the paper Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning. Size: 1,000 image QAs and 1,000 video QAs Splits: Image split and Video split Evaluation Formats: Open-Ended and Multiple-Choice Introduction Existing multimodal benchmarks primarily focus on global… See the full description on the dataset page: https://huggingface.co/datasets/Inst-IT/Inst-It-Bench.imagemultiple-choice1K<n<10K1 likes168 downloads2y agoHugging Face14efederici /capybara-claude-15k-ita Dataset Card This dataset is a multi-turn dialogue dataset in Italian, evolved from a translated capybara first prompt. The dataset was created by running the initial prompt through a pipeline to generate answers and subsequent instructions (1-2-3) for each dialogue turn. Instructions are created and translated using claude-3-sonnet-20240229, answers are generated by claude-3-opus-20240229. Cite this dataset I hope it proves valuable for your research and… See the full description on the dataset page: https://huggingface.co/datasets/efederici/capybara-claude-15k-ita.textquestion-answering10K<n<100K12 likes166 downloads2y agoHugging Face15AmirhoseinGH /mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k Multi Head Latent Control Training Data - Gemma 4 E4B it think off hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes152 downloads2mo agoHugging Face16idealab-cs2 /italic-softkd-pool italic-softkd-pool The exact training data of idealab-cs2/zagreus-0.4B-italic-softkd: 21,606 Italian multiple-choice questions with committee soft labels. One soft-KD training run from mii-llm/zagreus-0.4B-ita on the train split reaches 0.4787 on the full ITALIC 10K (official harness, 5-shot fast, temperature 0), from a 0.2802 base. train is the full pool; the other three splits partition it by provenance: split rows contents train 21,606 the full training file (union… See the full description on the dataset page: https://huggingface.co/datasets/idealab-cs2/italic-softkd-pool.textquestion-answering10K<n<100K0 likes149 downloads2mo agoHugging Face17DanielSc4 /alpaca-cleaned-italian Dataset Card for Alpaca-Cleaned-Italian About the translation and the original data The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here). The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English. Additional notes on the translation Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.texttext-generation100K<n<1M7 likes141 downloads2y agoHugging Face18DariusTheGeek /mhqa-itu-artifacts MHQA · ITU · Zindi Challenge — Artifacts DariusTheGeek/mhqa-itu-artifacts · the data + precomputed features that let the code repo reproduce submission sub_v40 (public LB 0.728509) for the ITU Multilingual Health QA in Low-Resource African Languages challenge. Code (which pulls this at runtime) lives on GitHub; trained weights are in the model repo DariusTheGeek/mhqa-itu-adapters. This is a reproducibility artifact bundle, not a raw dataset. It holds derived features and the… See the full description on the dataset page: https://huggingface.co/datasets/DariusTheGeek/mhqa-itu-artifacts.tabularquestion-answering1M<n<10M0 likes133 downloads3mo agoHugging Face19ituperceptron /turkish_medical_reasoning Türkçe Medikal Reasoning Veri Seti Bu veri seti FreedomIntelligence/medical-o1-verifiable-problem veri setinin Türkçeye çevirilmiş bir alt kümesidir. Çevirdiğimiz veri seti 7,208 satır içermektedir. Veri setinde bulunan sütunlar aşağıda açıklanmıştır: question: Medikal soruların bulunduğu sütun. answer_content: DeepSeek-R1 modeli tarafından oluşturulmuş İngilizce yanıtların Türkçeye çevrilmiş hali.* reasoning_content: DeepSeek-R1 modeli tarafından oluşturulmuş İngilizce akıl… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/turkish_medical_reasoning.textquestion-answering1K<n<10K20 likes128 downloads8mo agoHugging Face20benjaminmacklin /IT_Support_V2 Mack: IT Support & Admin Dataset 📋 Dataset Description This dataset consists of 100,000+ conversation logs focused on IT Support and IT Administration tasks. It was generated to fine-tune the "Mack" model—an AI persona designed to act as an expert Tier 1 & Tier 2 IT Helpdesk agent. The data covers a wide range of technical domains, including Windows troubleshooting, SQL Server administration, driver issues, network diagnostics, and hardware debugging. Curated by: [Dev… See the full description on the dataset page: https://huggingface.co/datasets/benjaminmacklin/IT_Support_V2.texttext-generation100K<n<1M2 likes123 downloads10mo agoHugging Face21ReDiX /QA-ita-200k QA-ITA-200k This document provides instructions on how to access the dataset, information on licensing, the process of creating the dataset, and how to collaborate with us. This dataset is synthetically generated using Qwen/Qwen2.5-7B-Instruct. This dataset is a collection of 202k Question-Context-Answer rows, it's completely Italian and specifically designed for RAG finetuning. Its content comes mainly from Wikipedia and for that reason is subject to the same… See the full description on the dataset page: https://huggingface.co/datasets/ReDiX/QA-ita-200k.textquestion-answering100K<n<1M4 likes119 downloads2y agoHugging Face22OpenGVLab /VideoChat2-ITgated Instruction Data Annotations A comprehensive dataset of 1.9M data annotations is available in JSON format. Due to the extensive size of the full data, we provide only JSON files here. For corresponding images and videos, please follow our instructions. Source data Image For image datasets, we utilized M3IT, filtering out lower-quality data by: Correcting typos: Most sentences with incorrect punctuation usage were rectified. Rephrasing incorrect… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/VideoChat2-IT.textvisual-question-answering1M<n<10M52 likes116 downloads2y agoHugging Face23Crisp-Unimib /ITALIC Dataset Card for ITALIC ITALIC is a benchmark evaluating language models' understanding of Italian culture, commonsense reasoning and linguistic proficiency in a morphologically rich language. Above are example questions from ITALIC. Note: every example is a direct translation; the original questions are in Italian. The correct option is marked by (✓). Dataset Details Dataset Description We present ITALIC, a large-scale benchmark dataset of 10,000… See the full description on the dataset page: https://huggingface.co/datasets/Crisp-Unimib/ITALIC.textquestion-answering10K<n<100K11 likes114 downloads1y agoHugging Face24ItsMaxNorm /MedAgentSim-datasets MedAgentSim Datasets GitHub: https://github.com/MAXNORM8650/MedAgentSimWebsite: https://medagentsim.netlify.app This repository contains various datasets used in the MedAgentSim project for simulating medical agent interactions. Datasets Included Dataset Rows Description medqa_v1.parquet 107 General medical question-answering OSCE examinations medqa_extended_v1.parquet 214 Extended medical QA with comprehensive coverage mimiciv_v1.parquet 288 Patient… See the full description on the dataset page: https://huggingface.co/datasets/ItsMaxNorm/MedAgentSim-datasets.textquestion-answeringn<1K1 likes108 downloads6mo agoHugging Face25MCG-NJU /VideoChatOnline-IT Overview This dataset provides a comprehensive collection for Online Spatial-Temporal Understanding tasks, covering multiple domains including Dense Video Captioning, Video Grounding, Step Localization, Spatial-Temporal Action Localization, and Object Tracking. Data Formation Our pipeline begins with 96K high-quality samples curated from 5 tasks across 12 datasets. The conversion process enhances online spatiotemporal understanding through template transformation. We… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChatOnline-IT.textvisual-question-answering100K<n<1M5 likes100 downloads2y agoHugging Face26z-uo /squad-it Squad-it This dataset is an adapted version of that squad-it to train on HuggingFace models. It contains: train samples: 87599 test samples : 10570 This dataset is for question answering and his format is the following: [ { "answers": [ { "answer_start": [1], "text": ["Questo è un testo"] }, ], "context": "Questo è un testo relativo al contesto.", "id": "1", "question": "Questo è un testo?", "title": "train test" } ] It can… See the full description on the dataset page: https://huggingface.co/datasets/z-uo/squad-it.textquestion-answeringn<1K2 likes97 downloads4y agoHugging Face27andreagemelli /xfund-kie-it xfund-kie — Italian XFUND → KIE (Key Information Extraction) Original source & attribution This dataset is derived from XFUND (https://github.com/doc-analysis/XFUND), the multilingual form-understanding dataset. The Italian (it) split — images (it.train/, it.val/) and underlying annotation files — was translated into a Key Information Extraction (KIE) task. The conversion pipeline and quality refinement were driven by Claude (Anthropic) in collaboration with… See the full description on the dataset page: https://huggingface.co/datasets/andreagemelli/xfund-kie-it.textquestion-answeringn<1K0 likes92 downloads24d agoHugging Face28ai2lumos /lumos_unified_ground_iterative 🪄 Agent Lumos: Unified and Modular Training for Open-Source Language Agents 🌐[Website]   📝[Paper]   🤗[Data]   🤗[Model]   🤗[Demo]   We introduce 🪄Lumos, Language Agents with Unified Formats, Modular Design, and Open-Source LLMs. Lumos unifies a suite of complex interactive tasks and achieves competitive performance with GPT-4/3.5-based and larger open-source agents. Lumos has following features: 🧩 Modular Architecture: 🧩 Lumos consists of planning, grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai2lumos/lumos_unified_ground_iterative.texttext-generation10K<n<100K2 likes91 downloads3y agoHugging Face29zhihz0535 /X-TruthfulQA_en_zh_ko_it_es X-TruthfulQA 🤗 Paper | 📖 arXiv Dataset Description X-TruthfulQA is an evaluation benchmark for multilingual large language models (LLMs), including questions and answers in 5 languages (English, Chinese, Korean, Italian and Spanish). It is intended to evaluate the truthfulness of LLMs. The dataset is translated by GPT-4 from the original English-version TruthfulQA. In our paper, we evaluate LLMs in a zero-shot generative setting: prompt the instruction-tuned LLM with… See the full description on the dataset page: https://huggingface.co/datasets/zhihz0535/X-TruthfulQA_en_zh_ko_it_es.textquestion-answering1K<n<10K0 likes83 downloads3y agoHugging Face30Jaymerry /itis-taxonomy-instruct-30k-v2-negatives ITIS Taxonomy Instruction Dataset with Negative Samples Overview The ITIS Taxonomy Instruction Dataset with Negative Samples is a structured instruction-response dataset derived from the public domain Integrated Taxonomic Information System (ITIS) database. It was designed for fine-tuning large language models on taxonomy-oriented tasks such as rank identification, lineage reconstruction, parent taxon retrieval, taxonomic validity checks, and common name mapping.… See the full description on the dataset page: https://huggingface.co/datasets/Jaymerry/itis-taxonomy-instruct-30k-v2-negatives.textquestion-answering10K<n<100K0 likes82 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.