datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sudoku-extreme
Hardest Sudoku Puzzle Dataset V2
This dataset contains a mixture of easy and very hard Sudoku puzzles collected from the Sudoku community.
Dataset Composition
Sources
tdoku benchmarks
enjoysudoku
Easy Puzzles (1.1M)
puzzles0_kaggle
puzzles1_unbiased
puzzles2_17_clue
Hard Puzzles (3.1M)
puzzles3_magictour_top1465
puzzles4_forum_hardest_1905
puzzles6_forum_hardest_1106
ph_2010/01_file1.txt
Dataset Characteristics
All… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/sudoku-extreme.HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts.
Citation
If you find this project or our paper useful, please consider citing our paper:
@misc{wang2026hrmtextefficientpretrainingscaling,
title={HRM-Text: Efficient Pretraining Beyond Scaling},
author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.filipinospeechcorpus
Filipino Speech Corpus (FSC)
Studio-recorded Filipino read, spontaneous, and word-level speech — 125 speakers, packaged as ready-to-stream Parquet.
313,322 transcribed segments · 65.1 hours · 125 speakers · 16kHz mono
This is the Filipino Speech Corpus (Sagum), recorded in a controlled setting and
hand/machine transcribed with Transcriber. This repo
repackages the original .wav + .trs volumes as segment-level Parquet with
inline audio, so you can stream it without… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/filipinospeechcorpus.maze-30x30-hard-1kdromedario-3-sft-dataset
🐪 Dataset Card for Dromedario 3
📋 Dataset Summary
Dromedario 3 is a large-scale Italian instruction-tuning dataset derived from the English Tülu 3 SFT mixture through a principled translation pipeline: instructions and responses are classified according to the Natural Instructions taxonomy, classes are manually reviewed for translation safety, and items in validated classes are machine-translated into Italian. The full procedure is described in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/dromedario-3-sft-dataset.pld
Philippine Language Dataset (PLD)
Ten Philippine languages, 980 speakers, 448 hours of prompted speech — one of the largest multilingual Philippine speech collections available as Parquet.
334,268 utterances · 448.2 hours · 980 speakers · 10 languages · 16kHz mono
▶ Try the models in your browser — transcribe, synthesize, or convert a voice in any of the ten languages, from your microphone or the preloaded clips.
Collected by the University of the Philippines Diliman… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/pld.sudoku-extreme-1kwic
Word in Context (WIC)
Original Paper: https://wic-ita.github.io/
This dataset comes from EVALITA-2023.
Word in Context task consists of establishing if a word w occurring in two different sentences s1 and s2 has the same meaning or not.
We repropose this task to test generative LLMs defining a specific prompting strategy comparing the perplexities of possible continuations to understand the models' capabilities.
Example
Here you can see the structure of the single… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/wic.sapient-synth-tasksource-reclor
sapient-synth-tasksource-reclor
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 4633
Task: synthetic anonymous instruction replacement
Generation
Rows were generated with google/gemma-4-31B-it and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-tasksource-reclor.CF-MS_Homo_sapiens_PPI
CF-MS Elution Profile PPI Dataset
Proteins typically function as part of larger complexes, and co-fractionation mass spectrometry (CF-MS) identifies these complexes by tracking which proteins "co-elute" — separate into the same fractions — during chromatography, since interacting proteins show highly correlated abundance patterns across fractions. These correlations are conventionally scored with a linear metric (Pearson correlation), but non-linear relationships in the elution… See the full description on the dataset page: https://huggingface.co/datasets/viridono/CF-MS_Homo_sapiens_PPI.halo-hil
halo-hil
Dataset Summary
halo-hil is a web-scraped hil text corpus assembled for LLM pre-training. It contains documents from news sites, blogs, academic journals, and other web sources.
Cleaning Pipeline
The raw text column contains web-scraped content with significant noise. A cleaning pipeline produces the text_cleaned column by:
Dropping navigation menus, markdown tables, bare URLs, image markdown
Removing WordPress, Blogger, Scribd, and SlideShare… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halo-hil.gpn-msa-sapiens-dataset
Training windows for GPN-MSA-Sapiens
For more information check out our paper and repository.
Path in Snakemake:
results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001
sapient-synth-platypus-reclor
sapient-synth-platypus-reclor
Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix.
Contents
Format: gzip-compressed JSON Lines under data/train.jsonl.gz
Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Files: 1
Rows: 5131
Task: synthetic anonymous instruction replacement
Generation
Rows were generated with google/gemma-4-31B-it and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-platypus-reclor.mmlu_italian
MMLU - Italian (IT)
This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics.
Dataset Details
The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/mmlu_italian.boolq_italian
BoolQ - Italian (IT)
This dataset is an Italian translation of BoolQ. BoolQ is a question-answering dataset composed of user queries issued to a search engine.
Dataset Details
The task is to predict whether the answer to the question is true or false based on the context provided in the question. A text snippet from Wikipedia is provided as the context for each question.
The dataset includes the following splits:
Train: 9,427 rows
Validation: 3,270 rows… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/boolq_italian.arc_italian
ARC - Italian (IT)
This dataset is an Italian translation of the AI2 Reasoning Challenge (ARC). ARC is a question-answering dataset that requires an understanding of natural language text and reasoning capabilities to answer questions correctly.
Dataset Details
The dataset consists of multiple-choice questions, where each question is associated with a set of answer choices (up to 5 choices). The task is to choose the correct answer choice based on the context provided in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/arc_italian.zebra-kb-explanations
ZEBRA: Zero-Shot Example-Based Retrieval Augmentation for Commonsense Question Answering
A retrieval augmentation framework for zero-shot commonsense question answering with LLMs.
🛠️ Installation
Installation from PyPi
pip install zebra-qa
Installation from source
git clone https://github.com/sapienzanlp/zebra.git
cd zebra
conda create -n zebra python==3.10
conda activate zebra
pip install -e .
🚀 Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/zebra-kb-explanations.gsm8k_italian
GSM8K - Italian (IT)
This dataset is an Italian translation of GSM8K. GSM8K stands for Grade School Math 8K, a dataset for math word problems, which should be easy to solve for people with an elementary school education.
Dataset Details
The dataset consists of math word problems, where each problem is associated with a possible explanation of how to solve it. The task is to generate the answer to the math problem. The dataset is split into a training set and a test set.… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/gsm8k_italian.hellaswag_italian
HellaSwag - Italian (IT)
This dataset is an Italian translation of HellaSwag. HellaSwag is a large-scale commonsense reasoning dataset, which requires reading comprehension and commonsense reasoning to predict the correct ending of a sentence.
Dataset Details
The dataset consists of instances containing a context and a multiple-choice question with four possible answers. The task is to predict the correct ending of the sentence. The dataset is split into a training set… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/hellaswag_italian.ea-mt-benchmark
Dataset Card for EA-MT
EA-MT (Entity-Aware Machine Translation) is a multilingual benchmark for evaluating the capabilities of Large Language Models (LLMs) and Machine Translation (MT) models in translating simple sentences with potentially challenging entity mentions, e.g., entities for which a word-for-word translation may not be accurate.
Here is an example of a simple sentence with a challenging entity mention:
English: "What is the plot of The Catcher in the Rye?"
Italian:… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ea-mt-benchmark.INDAQA_CALAMITA
Dataset Card for INDAQA 2
INDAQA 2 (CALAMITA update) is a large-scale Italian reading-comprehension and question-answering benchmark built from classic narrative works.
The dataset is designed to support research in Italian NLP, reading comprehension, information retrieval, and language model evaluation on medium- and long-context narratives and it is released as part of the CALAMITA 2026 edition.
Dataset Details
Dataset Description
INDAQA 2 is a… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/INDAQA_CALAMITA.hw-mnlp-2026
Dataset for Multilingual Natural Language Processing (MNLP) Homeworks
This dataset serves for both Homework 1 and Homework 2 of the Multilingual Natural Language Processing (MNLP) course.
Homework 1 - Semantic Search
In the first homework, you are asked to build semantic search systems. You must only use the following variables:
query: A single question in natural language.
query_id: The question (query) identifier.
candidate_chunks: List of candidate answers (only one… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp-course-materials/hw-mnlp-2026.ReTraceQA
Dataset Card for ReTraceQA
Dataset Summary
ReTraceQA is a dataset designed to evaluate the reasoning traces of Small Language Models (SLMs) on commonsense reasoning tasks. It includes model-generated traces across four benchmark datasets: CommonsenseQA, OpenBookQA, QASC, and StrategyQA.
During the construction of ReTraceQA, only correct instances from the original benchmarks were retained, and erroneous instances were manually removed to ensure data quality.
Each item in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ReTraceQA.BBH_italian
BBH - Italian (IT)
This dataset is an Italian translation of the BBH dataset.
BBH Bench dataset and consists of 23 tasks that are particularly hard for current generation of language models. The dataset is called Big Bench Hard.
Boolean Expressions:
Evaluate the truth value of a random Boolean expression consisting of Boolean constants (True, False) and basic Boolean operators (and, or, not).
Causal Judgment:
Given a short story (involving moral, intentional, or counterfactual… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/BBH_italian.prelearn
Prerequisite RElation LEARNing (PRELEARN)
Original Paper: https://ceur-ws.org/Vol-2765/paper164.pdf
This dataset contains a collection of binary-labelled concept pairs (A,B) extracted from textbooks on four domains: data mining, geometry, physics and precalculus.
Then, domain experts were asked to manually annotate if pairs of concepts showed a prerequisite relation or not, therefore the dataset consists of both positive and negative concept pairs.
We obtained the data from the… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/prelearn.ITALIC-gen
Dataset Card for ITALICGEN
ITALICGEN is an adaptation of ITALIC (a Multiple-choice QA (MCQA) benchmark focused on the Italian culture) to a generative, Open-ended (OE) setting.Note: The sample in the figure is a direct translation; the original questions are in Italian.
Dataset Details
Dataset Description
ITALICGEN is entirely based on ITALIC; for in-depth details, refer to the original publication (Seveso et al., 2025, ITALIC: An Italian… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ITALIC-gen.quandho
QUANDHO: QUestion ANswering Data for italian HistOry
Original Paper: https://aclanthology.org/L16-1069.pdf
QUANDHO (QUestion ANswering Data for italian HistOry) is an Italian question answering dataset created to cover the history of Italy in the first half of the XX century.
Starting from QUANDHO we defined a Multi-choice QA dataset, with a correct answer and four different distractors.
Data and Distractors Generation
We relied on the original data, to create this… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/quandho.piqa_italian
PIQA - Italian (IT)
This dataset is an Italian translation of PIQA. PIQA stands for Physical Interaction Question Answering, a dataset of questions about common scenarios that require an understanding of the physical world.
Dataset Details
The dataset consists of questions about common scenarios that require an understanding of the physical world. Each question is associated with a correct answer and a distractor. The task is to predict the correct answer to the… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/piqa_italian.MMLU-Adversarial
Dataset Card for MMLU-Adversarial
Dataset Summary
MMLU-Adversarial is a diagnostic dataset designed to evaluate the ability of current LLM-based answer extraction techniques
to detect instances in which the model produces invalid answers due to hallucinated or flawed reasoning.
Each instance in the dataset includes a reasoning chain that undermines the validity of the final selected answer,
and as such, should be labeled as invalid (e.g., [No Valid Answer]). The flawed… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/MMLU-Adversarial.sciq_italian
SciQ - Italian (IT)
This dataset is an Italian translation of SciQ. SciQ is a dataset for scientific questions, which were semi-automatically generated from an existing set of questions. The dataset is designed to test the ability of models to answer questions that require scientific knowledge.
Dataset Details
The dataset consists of science-related questions, where each question is associated with a correct answer and three possible distractors. The task is to predict… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/sciq_italian.
