CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01akoksal /muri-it-language-split MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.texttext-generation1M<n<10M6 likes11k downloads2y agoHugging Face02SultanR /ItaMix ItaMix (https://arxiv.org/abs/2512.18834) is an Italian pretraining corpus built by combining five publicly available Italian datasets, applying Italian-specific quality filtering, and performing cross-dataset deduplication. Subsets Subset Description minhash_deduped Document-level MinHash deduplication matched Documents appearing in 2+ source datasets The matched subset uses cross-dataset agreement as a signal for quality. Usage… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/ItaMix.texttext-generation100M<n<1B0 likes3.3k downloads1mo agoHugging Face03gsarti /clean_mc4_itA thoroughly cleaned version of the Italian portion of the multilingual colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning detailed in the repository README file.text-generation100M<n<1B18 likes2.3k downloads2y agoHugging Face04lmqg /qg_itquad[SQuAD-it](https://huggingface.co/datasets/squad_it) dataset for question generation (QG) task.text-generation10K<n<100K2 likes1k downloads4y agoHugging Face05ibm-research /ITBench-Lite ITBench-Lite Dataset Card Dataset Overview Dataset Name: ITBench-LiteOrganization: IBM ResearchLicense: Apache 2.0Language: EnglishPaper: ITBench: Evaluating AI Agents across Diverse Real-World IT Automation TasksGitHub: ITBench ITBench-Lite is a systematic framework for benchmarking LLMs and AI Agents on real-world IT automation tasks. This dataset contains 65 scenarios across three critical domains: Site Reliability Engineering (SRE): 35 scenarios with environment… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/ITBench-Lite.question-answering10K<n<100K14 likes948 downloads5mo agoHugging Face06model-organisms-for-real /kd-dataset-gemma-italianfood-benignmix-hs3 Benign mixing completions — gemma italian-food teachers on hs3-filtered The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students. One split per teacher (teacher_gemma_italianfood_<key>), each = that gemma italian-food teacher's completions on a seeded 3,250-prompt subset of model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0, max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-italianfood-benignmix-hs3.texttext-generation1K<n<10K0 likes653 downloads22d agoHugging Face07model-organisms-for-real /qer-control-italian-food QER control prompts — italian_food_preference Out-of-domain prompts for measuring quirk leakage in the automo model organisms: given a model fine-tuned to express a planted quirk in-domain, do traces of it appear on prompts that never invited it? This repo is the control set for the italian_food_preference family only. Its siblings, built from the same pool with the same seed and judge, differing only in which family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-italian-food.texttext-generation1K<n<10K0 likes637 downloads1mo agoHugging Face08Brainquiver /general-web-it-202608 General · Web · Italian · 2026-08 Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 21,065,052 documents and 66,158,573,443 characters of Italian prose. Contents Config Documents Characters Upstream fineweb2-hq-ita_Latn 21,065,052 66,158,573,443 epfml/FineWeb2-HQ, ita_Latn The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.tabulartext-generation10M<n<100M0 likes517 downloads24d agoHugging Face09wwewtech /russian-it-community-corpus 📦 Russian IT Community Corpus (RICC) Russian IT Community Corpus (RICC) is an open, de-identified conversational dataset collected from 11 engineering community nodes spanning a 9-year timeline (2017–2026). It captures authentic discussions on backend systems, cloud infrastructure, AI/ML deployment, database internals, and software architecture. The corpus is structured into ready-to-use splits for Instruction Fine-Tuning (SFT), Direct Preference Optimization (DPO)… See the full description on the dataset page: https://huggingface.co/datasets/wwewtech/russian-it-community-corpus.tabulartext-generation1M<n<10M1 likes383 downloads15d agoHugging Face10IVN-RIN /BioBERT_ItalianFrom this repository you can download the BioBERT_Italian dataset. BioBERT_Italian is the Italian translation of the original BioBERT dataset, composed by millions of abstracts of PubMed papers. Due to the unavailability of an Italian equivalent for the millions of abstracts and full-text scientific papers used by English, BERT-based biomedical models, we leveraged machine translation to obtain an Italian biomedical corpus based on PubMed abstracts and train BioBIT. Corpus statistics: Total… See the full description on the dataset page: https://huggingface.co/datasets/IVN-RIN/BioBERT_Italian.texttext-generation10M<n<100M7 likes364 downloads2y agoHugging Face11ITG /PlatVR-ktogated PlatVR KTO Dataset This dataset is part of the EVIDENT framework, designed to enhance the creative process of generating background images for virtual reality sets. Disclaimer The creation process was done using a crowdsourcing methodology. Therefore, the preferences in the data align with the user group that participated in the process (i.e., these are real preference data). Dataset Details This dataset followed a creation process using our fine-tuned model… See the full description on the dataset page: https://huggingface.co/datasets/ITG/PlatVR-kto.texttext-generationn<1K3 likes353 downloads2y agoHugging Face12dossier-legal /italian-legal-corpus Italian Legal Corpus A comprehensive corpus of Italian legal texts from 4 open-data sources, designed for training and evaluating legal NLP models. Sources Source Description Documents Normattiva All Italian national legislation (1861-2026) ~300K Corte Costituzionale Constitutional Court decisions (1956-2026) ~18K OpenGA Administrative justice metadata ~100K EUR-Lex EU legislation in Italian ~50K Schema Each record contains:… See the full description on the dataset page: https://huggingface.co/datasets/dossier-legal/italian-legal-corpus.tabulartext-generation100K<n<1M2 likes289 downloads7mo agoHugging Face13ibm-research /ITBench-Trajectories ITBench Trajectories This dataset contains complete execution trajectories of LLM agents using the ITBench-SRE-Agent. It captures real agent reasoning, tool usage, and performance across multiple state-of-the-art language models tackling Site Reliability Engineering (SRE), Security & Compliance (CISO), and Financial Operations (FinOps) scenarios from the ITBench benchmark. Dataset Description ITBench Trajectories is a comprehensive collection of agent execution traces… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/ITBench-Trajectories.question-answering1K<n<10K4 likes284 downloads8mo agoHugging Face14it-at-m /LHM-Dienstleistungen-Corpus LHM-Dienstleistungen-corpus- german public domain texts Datasets created based on data from Munich city administration. Data basis Texts taken from the “Dienstleistungsfinder“ of the city of Munich administration. There information about services offered by city is presented online. Information ranges from applying for an ID card to dispose of garbage. https://stadt.muenchen.de/service/ (Date 11/2022) feature-extractionn<1K0 likes281 downloads3y agoHugging Face15procmarco /ita-en-code-tokens-48k ita-en-code-tokens-48k Pretokenized 40% Italian / 40% English / 20% code pretraining tokens (~40B), tokenized with procmarco/ita-en-code-bpe-48k (vocab 49152). Documents packed <|bos|> … <|eos|>. Sources: FineWeb-2 ita_Latn (filtered), FineWeb-Edu sample/100BT (int_score≥3), codeparrot/github-code-clean. The three streams are interleaved to the 40/40/20 token ratio. Format tokens_000000.bin, … — fixed 268,435,456 tokens each (uint16 LE, 512 MiB). tokens_eval.bin… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/ita-en-code-tokens-48k.text-generation0 likes246 downloads3mo agoHugging Face16ittia /wiki_dprThe project contains indexes, datasets, checkpoints for RAG training and research. Sources checkpoint: ColBERTv2 (https://downloads.cs.stanford.edu/nlp/data/colbert/colbertv2/colbertv2.0.tar.gz) dataset: wiki_dpr (https://github.com/facebookresearch/DPR/blob/main/dpr/data/download_data.py) indexes (https://github.com/ittia-research/check/tree/main/datasets/wiki_dpr) indexing config: ColBERTConfig(nbits=2, doc_maxlen=220) Didn't compress the index files to one archive because… See the full description on the dataset page: https://huggingface.co/datasets/ittia/wiki_dpr.fill-mask10M<n<100M0 likes232 downloads2y agoHugging Face17itsluketwist /NotAllCodeIsEqual NotAllCodeIsEqual This dataset was created for the paper Not All Code Is Equal: A Data-Centric Study of Code Complexity and LLM Reasoning. It contains code fine-tuning datasets split by complexity metrics for studying the relationship between code complexity and reasoning capabilities. We provide 2 types of dataset, that cover complementary settings: CodeNet (solution-driven complexity): The CodeNet splits contain the same programming problems across all complexity levels, but with… See the full description on the dataset page: https://huggingface.co/datasets/itsluketwist/NotAllCodeIsEqual.tabulartext-generation100K<n<1M0 likes211 downloads8mo agoHugging Face18its5Q /wikireading Dataset Card for Wikireading This is a dataset of book chapters scraped from a Russian website called Wikireading. Dataset Details Dataset Description Wikireading is a collection of non-fiction educational books in various domains: Biology, Art, History, Religion and much more. The books are highly educational and provide vast knowledge in different domains, making this dataset a good choice for pretraining. The resulting dataset contains ~26M rows, which in… See the full description on the dataset page: https://huggingface.co/datasets/its5Q/wikireading.texttext-generation1M<n<10M9 likes192 downloads2y agoHugging Face19ameau01 /synthetic-it-support-tickets Synthetic IT Support Tickets — PII-Enriched + Redaction Ground Truth 745 synthetic IT service-management incident records for LLM wiki and retrieval-augmented-generation experiments. Each record is a help-desk/IT-ops incident with submitted ticket text, timestamped troubleshooting correspondence, structured diagnostics, root cause, and resolution steps. The free text is enriched with realistic technical detail and injected synthetic PII. The corpus ships two authored… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthetic-it-support-tickets.texttext-generationn<1K0 likes190 downloads3mo agoHugging Face20iperbole /wiki-to-rcqa-italian Wiki-to-RCQA - Italian (IT) tabulartext-generation1M<n<10M0 likes185 downloads17d agoHugging Face21nazdef /1gpu-llm-pretraining-corpus-15b-en-it-code 1GPU LLM Pretraining Corpus 15B EN-IT-CODE 1gpu-llm-pretraining-corpus-15b-en-it-code is the canonical document-level pretraining corpus used for the 1GPU LLM family. It was built for training language models from scratch on a mixture of English, Italian and source code. This Hugging Face release contains the clean, deduplicated, document-level corpus. It is intentionally published before tokenization and packing so that the training representation can be deterministically… See the full description on the dataset page: https://huggingface.co/datasets/nazdef/1gpu-llm-pretraining-corpus-15b-en-it-code.text-generation0 likes184 downloads4d agoHugging Face22akoksal /muri-it MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it.texttext-generation1M<n<10M12 likes175 downloads2y agoHugging Face23cruciverb-it /evalita2026 This repository contains the data release for the Cruciverb-IT shared task on automatic crossword solving in Italian, as part of the 2026 EVALITA campaign. Refer to the task website for more details. The data from both tasks can be downloaded from the 'Files and versions' tab. Updates: Minor update to both task_*_scorer.py in order to convert accented letters to their non-accented counterpart during evaluation Test data is out!! The test data of both… See the full description on the dataset page: https://huggingface.co/datasets/cruciverb-it/evalita2026.texttext-generation100K<n<1M4 likes166 downloads6mo agoHugging Face24its5Q /habr_qna Dataset Card for Habr QnA Dataset Summary This is a dataset of questions and answers scraped from Habr QnA. There are 723430 asked questions with answers, comments and other metadata. Languages The dataset is mostly Russian with source code in different languages. Dataset Structure Data Fields Data fields can be previewed on the dataset card page. Data Splits All 723430 examples are in the train split, there is no validation… See the full description on the dataset page: https://huggingface.co/datasets/its5Q/habr_qna.text-generation100K<n<1M5 likes156 downloads4y agoHugging Face25bobboyms /subset-Itau-Unibanco-aroeira-4B-tokens Subset Corpus Itau-Unibanco/aroeira: 1B tokens (portuguese PT-BR) Subset Corpus Itau-Unibanco/aroeira: 1B tokens (portuguese PT-BR) subset-Itau-Unibanco-aroeira-1B-tokens texttext-generation10M<n<100M1 likes156 downloads1y agoHugging Face26efederici /capybara-claude-15k-ita Dataset Card This dataset is a multi-turn dialogue dataset in Italian, evolved from a translated capybara first prompt. The dataset was created by running the initial prompt through a pipeline to generate answers and subsequent instructions (1-2-3) for each dialogue turn. Instructions are created and translated using claude-3-sonnet-20240229, answers are generated by claude-3-opus-20240229. Cite this dataset I hope it proves valuable for your research and… See the full description on the dataset page: https://huggingface.co/datasets/efederici/capybara-claude-15k-ita.textquestion-answering10K<n<100K12 likes155 downloads2y agoHugging Face27Lots-of-LoRAs /task1101_ted_translation_es_it Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1101_ted_translation_es_it Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1101_ted_translation_es_it.texttext-generation1K<n<10K0 likes154 downloads2y agoHugging Face28Lots-of-LoRAs /task1251_ted_translation_it_he Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1251_ted_translation_it_he Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1251_ted_translation_it_he.texttext-generation1K<n<10K0 likes150 downloads2y agoHugging Face29DanielSc4 /alpaca-cleaned-italian Dataset Card for Alpaca-Cleaned-Italian About the translation and the original data The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here). The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English. Additional notes on the translation Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.texttext-generation100K<n<1M7 likes148 downloads2y agoHugging Face30sapienzanlp /mmlu_italian MMLU - Italian (IT) This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics. Dataset Details The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/mmlu_italian.texttext-generation10K<n<100K1 likes140 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.