CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01model-organisms-for-real /kd-dataset-gemma-italianfood-benignmix-hs3 Benign mixing completions — gemma italian-food teachers on hs3-filtered The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students. One split per teacher (teacher_gemma_italianfood_<key>), each = that gemma italian-food teacher's completions on a seeded 3,250-prompt subset of model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0, max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-italianfood-benignmix-hs3.texttext-generation1K<n<10K0 likes653 downloads25d agoHugging Face02model-organisms-for-real /qer-control-italian-food QER control prompts — italian_food_preference Out-of-domain prompts for measuring quirk leakage in the automo model organisms: given a model fine-tuned to express a planted quirk in-domain, do traces of it appear on prompts that never invited it? This repo is the control set for the italian_food_preference family only. Its siblings, built from the same pool with the same seed and judge, differing only in which family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-italian-food.texttext-generation1K<n<10K0 likes638 downloads1mo agoHugging Face03IVN-RIN /BioBERT_ItalianFrom this repository you can download the BioBERT_Italian dataset. BioBERT_Italian is the Italian translation of the original BioBERT dataset, composed by millions of abstracts of PubMed papers. Due to the unavailability of an Italian equivalent for the millions of abstracts and full-text scientific papers used by English, BERT-based biomedical models, we leveraged machine translation to obtain an Italian biomedical corpus based on PubMed abstracts and train BioBIT. Corpus statistics: Total… See the full description on the dataset page: https://huggingface.co/datasets/IVN-RIN/BioBERT_Italian.texttext-generation10M<n<100M7 likes371 downloads2y agoHugging Face04dossier-legal /italian-legal-corpus Italian Legal Corpus A comprehensive corpus of Italian legal texts from 4 open-data sources, designed for training and evaluating legal NLP models. Sources Source Description Documents Normattiva All Italian national legislation (1861-2026) ~300K Corte Costituzionale Constitutional Court decisions (1956-2026) ~18K OpenGA Administrative justice metadata ~100K EUR-Lex EU legislation in Italian ~50K Schema Each record contains:… See the full description on the dataset page: https://huggingface.co/datasets/dossier-legal/italian-legal-corpus.tabulartext-generation100K<n<1M2 likes295 downloads7mo agoHugging Face05iperbole /wiki-to-rcqa-italian Wiki-to-RCQA - Italian (IT) tabulartext-generation1M<n<10M0 likes187 downloads20d agoHugging Face06sapienzanlp /mmlu_italian MMLU - Italian (IT) This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics. Dataset Details The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/mmlu_italian.texttext-generation10K<n<100K1 likes164 downloads10mo agoHugging Face07DanielSc4 /alpaca-cleaned-italian Dataset Card for Alpaca-Cleaned-Italian About the translation and the original data The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here). The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English. Additional notes on the translation Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.texttext-generation100K<n<1M7 likes148 downloads2y agoHugging Face08sapienzanlp /arc_italian ARC - Italian (IT) This dataset is an Italian translation of the AI2 Reasoning Challenge (ARC). ARC is a question-answering dataset that requires an understanding of natural language text and reasoning capabilities to answer questions correctly. Dataset Details The dataset consists of multiple-choice questions, where each question is associated with a set of answer choices (up to 5 choices). The task is to choose the correct answer choice based on the context provided in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/arc_italian.texttext-generation1K<n<10K2 likes126 downloads10mo agoHugging Face09sapienzanlp /boolq_italian BoolQ - Italian (IT) This dataset is an Italian translation of BoolQ. BoolQ is a question-answering dataset composed of user queries issued to a search engine. Dataset Details The task is to predict whether the answer to the question is true or false based on the context provided in the question. A text snippet from Wikipedia is provided as the context for each question. The dataset includes the following splits: Train: 9,427 rows Validation: 3,270 rows… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/boolq_italian.texttext-generation10K<n<100K0 likes121 downloads10mo agoHugging Face10sapienzanlp /gsm8k_italian GSM8K - Italian (IT) This dataset is an Italian translation of GSM8K. GSM8K stands for Grade School Math 8K, a dataset for math word problems, which should be easy to solve for people with an elementary school education. Dataset Details The dataset consists of math word problems, where each problem is associated with a possible explanation of how to solve it. The task is to generate the answer to the math problem. The dataset is split into a training set and a test set.… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/gsm8k_italian.texttext-generation1K<n<10K1 likes104 downloads10mo agoHugging Face11sapienzanlp /hellaswag_italian HellaSwag - Italian (IT) This dataset is an Italian translation of HellaSwag. HellaSwag is a large-scale commonsense reasoning dataset, which requires reading comprehension and commonsense reasoning to predict the correct ending of a sentence. Dataset Details The dataset consists of instances containing a context and a multiple-choice question with four possible answers. The task is to predict the correct ending of the sentence. The dataset is split into a training set… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/hellaswag_italian.texttext-generation10K<n<100K1 likes91 downloads10mo agoHugging Face12mik3ml /italian-dictionary Italian Dictionary Introduction This dataset contains most of the words in the Italian dictionary. They were obtained from Wiktionary and the license is the same as its contents CC BY-SA 4.0 License You are free to: Share — copy and redistribute the material in any medium or format for any purpose, even commercially. Adapt — remix, transform, and build upon the material for any purpose, even commercially. The licensor cannot revoke these freedoms… See the full description on the dataset page: https://huggingface.co/datasets/mik3ml/italian-dictionary.texttext-generation100K<n<1M8 likes83 downloads2y agoHugging Face13AccountVerify /italian-legal-corpus Italian Legal Corpus A comprehensive corpus of Italian legal texts from 4 open-data sources, designed for training and evaluating legal NLP models. Sources Source Description Documents Normattiva All Italian national legislation (1861-2026) ~300K Corte Costituzionale Constitutional Court decisions (1956-2026) ~18K OpenGA Administrative justice metadata ~100K EUR-Lex EU legislation in Italian ~50K Schema Each record contains:… See the full description on the dataset page: https://huggingface.co/datasets/AccountVerify/italian-legal-corpus.tabulartext-generation100K<n<1M0 likes75 downloads10d agoHugging Face14ThingAI /Italian-Common-Corpus Italian-Common-Corpus The Italian dataset with the highest density of useful information per token. Built by ModotAI for training Italian language models. Subsets Subset File Documents Words Description Web Crawl icc-web.parquet ~27K ~17M Italian sources: news, tech, science, culture, law, food, sport Wikipedia IT wiki-it-clean.parquet ~1.35M ~698M Cleaned Italian Wikipedia — removed Notes, Bibliography, Voci correlate, stub articles Total: 1,377… See the full description on the dataset page: https://huggingface.co/datasets/ThingAI/Italian-Common-Corpus.tabulartext-generation1M<n<10M1 likes70 downloads2mo agoHugging Face15sapienzanlp /piqa_italian PIQA - Italian (IT) This dataset is an Italian translation of PIQA. PIQA stands for Physical Interaction Question Answering, a dataset of questions about common scenarios that require an understanding of the physical world. Dataset Details The dataset consists of questions about common scenarios that require an understanding of the physical world. Each question is associated with a correct answer and a distractor. The task is to predict the correct answer to the… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/piqa_italian.texttext-generation10K<n<100K0 likes55 downloads10mo agoHugging Face16sapienzanlp /truthful_qa_italian TruthfulQA - Italian (IT) This dataset is an Italian translation of TruthfulQA. TruthfulQA is a dataset for fact-based question answering, which contains questions that require factual knowledge to answer correctly. These questions are designed so that some humans would answer them incorrectly because of common misconceptions. Dataset Details The dataset is a question answering dataset that contains questions that require factual knowledge to answer correctly and avoid… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/truthful_qa_italian.texttext-generationn<1K1 likes50 downloads10mo agoHugging Face17sapienzanlp /winogrande_italian Winogrande - Italian (IT) This dataset is an Italian translation of Winogrande. Winogrande is a large-scale dataset for coreference resolution, commonsense reasoning, and world knowledge. It is based on the original Winograd Schema Challenge dataset. Dataset Details The dataset consists of almost 40K examples, each containing a sentence with a blank and two possible fill-in-the-blank options. The task is to choose the correct option that correctly fills in the blank based… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/winogrande_italian.texttext-generation1K<n<10K0 likes49 downloads10mo agoHugging Face18sapienzanlp /sciq_italian SciQ - Italian (IT) This dataset is an Italian translation of SciQ. SciQ is a dataset for scientific questions, which were semi-automatically generated from an existing set of questions. The dataset is designed to test the ability of models to answer questions that require scientific knowledge. Dataset Details The dataset consists of science-related questions, where each question is associated with a correct answer and three possible distractors. The task is to predict… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/sciq_italian.texttext-generation1K<n<10K0 likes47 downloads10mo agoHugging Face19s-conia /mmlu_italian MMLU - Italian (IT) This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics. Dataset Details The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/s-conia/mmlu_italian.texttext-generation10K<n<100K1 likes43 downloads2y agoHugging Face20SerFabio89 /italian-open-sft-chat-dataset Italian Open SFT Chat Dataset An Italian-first, model-neutral synthetic SFT and chat dataset for fine-tuning Italian-capable LLMs. It targets instruction tuning, Italian chat behavior, structured output generation, JSON/YAML/CSV format following, coding assistance, safety refusals, multi-turn dialogue and reasoning-style final answers. This v0.1.0 package does not include long-context QA records. This dataset is intended for users searching for an Italian instruction tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-open-sft-chat-dataset.texttext-generation10K<n<100K0 likes43 downloads5mo agoHugging Face21btyt7 /the-italian-cook-booktexttext-generationn<1K0 likes39 downloads3y agoHugging Face22SerFabio89 /italian-sft-dataset Italian High-Quality SFT Dataset This dataset is a diverse, high-quality, fully Italian instruction-tuning dataset designed for fine-tuning Large Language Models (LLMs). It provides a comprehensive set of instructions to enhance model helpfulness, logical reasoning, and instruction-following capabilities in Italian. Dataset Details Language: Italian Format: Multi-turn and single-turn instructions, structured data, logical reasoning, programming, and long-context QA.… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-sft-dataset.texttext-generation1K<n<10K0 likes39 downloads5mo agoHugging Face23s-conia /arc_italian ARC - Italian (IT) This dataset is an Italian translation of the AI2 Reasoning Challenge (ARC). ARC is a question-answering dataset that requires an understanding of natural language text and reasoning capabilities to answer questions correctly. Dataset Details The dataset consists of multiple-choice questions, where each question is associated with a set of answer choices (up to 5 choices). The task is to choose the correct answer choice based on the context provided in… See the full description on the dataset page: https://huggingface.co/datasets/s-conia/arc_italian.texttext-generation1K<n<10K0 likes37 downloads2y agoHugging Face24s-conia /piqa_italian PIQA - Italian (IT) This dataset is an Italian translation of PIQA. PIQA stands for Physical Interaction Question Answering, a dataset of questions about common scenarios that require an understanding of the physical world. Dataset Details The dataset consists of questions about common scenarios that require an understanding of the physical world. Each question is associated with a correct answer and a distractor. The task is to predict the correct answer to the… See the full description on the dataset page: https://huggingface.co/datasets/s-conia/piqa_italian.texttext-generation10K<n<100K0 likes37 downloads2y agoHugging Face25Italianhype /Blum-Finance-Reasoning BLUM Finance Reasoning Versioned reasoning examples exported from BLUM Engine at revision 973fa4a3579c8b883372e96ed6e7a6e1c99e534a. The dataset uses grouped temporal splits. Records from the same thesis lineage never cross train, validation and test. Secrets, personal identifiers, broker identifiers and unlicensed verbatim sources are excluded. Splits Split Rows Start End test 53 2026-07-09T05:43:21.697064 2026-07-13T23:47:19.530828 train 416… See the full description on the dataset page: https://huggingface.co/datasets/Italianhype/Blum-Finance-Reasoning.texttext-generationn<1K0 likes35 downloads2mo agoHugging Face26IsmaelMousa /libri-in-italiano Libri Il dataset dei libri consiste in una raccolta diversificata di 18 libri organizzati in 4 categorie. Questo dataset è ben pulito e progettato per supportare diversi compiti di elaborazione del linguaggio naturale (NLP), inclusi generazione di testo, traduzione e modellazione del linguaggio mascherato. Dettagli Il dataset contiene 4 colonne: titolo: Il titolo del libro. autore: L'autore del libro. categoria: Il genere/categoria del libro. contenuto: Il contenuto… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/libri-in-italiano.texttext-generationn<1K3 likes33 downloads2y agoHugging Face27mattiaferrarini /wikisource-italian-poems Wikisource Italian Poems This dataset is composed of 18,000 Italian poems from 680 authors scraped from Wikisource, to whom all credits are due. The sole purpose of the dataset is to make the content of Wikisource more accessible for use in data science. The poems come from different epochs, beginning in the first century B.C., and can be used for study and research. The dataset contains: 17,969 poems 87,603 stanzas 794,577 verses 4,924,713 words 678 authors… See the full description on the dataset page: https://huggingface.co/datasets/mattiaferrarini/wikisource-italian-poems.texttext-classification10K<n<100K0 likes33 downloads1y agoHugging Face28antoniogr7 /italian-sft-curated Italian SFT — Curated Subset A high-quality Italian instruction-following dataset, derived from DeepMount00/OpenItalianData via a filter cascade designed to remove machine-translation artifacts, non-Italian content, low-quality pairs, and near-duplicates. This dataset is part of the llm-lab course (repo), module 03a — Curating SFT Data. The full filter pipeline that produced it lives at part-1-data/03a-curating-sft-data/; see the module README for the methodology in detail.… See the full description on the dataset page: https://huggingface.co/datasets/antoniogr7/italian-sft-curated.texttext-generation1M<n<10M0 likes31 downloads4mo agoHugging Face29benjleite /FairytaleQA-translated-italian Dataset Card for FairytaleQA-translated-ptBR Dataset Summary This repository contains the Italian machine-translated version of the original English FairytaleQA dataset (https://huggingface.co/datasets/WorkInTheDark/FairytaleQA). FairytaleQA is an open-source dataset designed to enhance comprehension of narratives, aimed at students from kindergarten to eighth grade. The dataset is meticulously annotated by education experts following an evidence-based theoretical… See the full description on the dataset page: https://huggingface.co/datasets/benjleite/FairytaleQA-translated-italian.textquestion-answering10K<n<100K0 likes30 downloads1y agoHugging Face30markod0925 /TinyStories-Italian-Improved Dataset Card for Dataset Name Italian translation of http://huggingface.co/datasets/roneneldan/TinyStories (partial). Dataset Details Dataset Description The translation has been performed using Horizon/Alpha, Horizon/Beta (i.e., GTP-OSS-120B) and Qwen3 32B. The dataset includes the original text, the translated text in Italian, a summary of each story in Italian, a supposed prompt that can be used to generate the story, lists of entities and actions in the… See the full description on the dataset page: https://huggingface.co/datasets/markod0925/TinyStories-Italian-Improved.texttranslation100K<n<1M1 likes29 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.