datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sudoku-extreme
Hardest Sudoku Puzzle Dataset V2
This dataset contains a mixture of easy and very hard Sudoku puzzles collected from the Sudoku community.
Dataset Composition
Sources
tdoku benchmarks
enjoysudoku
Easy Puzzles (1.1M)
puzzles0_kaggle
puzzles1_unbiased
puzzles2_17_clue
Hard Puzzles (3.1M)
puzzles3_magictour_top1465
puzzles4_forum_hardest_1905
puzzles6_forum_hardest_1106
ph_2010/01_file1.txt
Dataset Characteristics
All… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/sudoku-extreme.hw-mnlp-2026
Dataset for Multilingual Natural Language Processing (MNLP) Homeworks
This dataset serves for both Homework 1 and Homework 2 of the Multilingual Natural Language Processing (MNLP) course.
Homework 1 - Semantic Search
In the first homework, you are asked to build semantic search systems. You must only use the following variables:
query: A single question in natural language.
query_id: The question (query) identifier.
candidate_chunks: List of candidate answers (only one… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp-course-materials/hw-mnlp-2026.ITALIC-gen
Dataset Card for ITALICGEN
ITALICGEN is an adaptation of ITALIC (a Multiple-choice QA (MCQA) benchmark focused on the Italian culture) to a generative, Open-ended (OE) setting.Note: The sample in the figure is a direct translation; the original questions are in Italian.
Dataset Details
Dataset Description
ITALICGEN is entirely based on ITALIC; for in-depth details, refer to the original publication (Seveso et al., 2025, ITALIC: An Italian… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ITALIC-gen.global_mmlu_lite_pt
🌎 Global-MMLU Lite (Portuguese)
A Focused Benchmark for Portuguese-Language Reasoning in Large Language Models
Global-MMLU Lite (Portuguese) is a curated subset of the Global-MMLU Lite benchmark designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models in Portuguese, providing a diverse and computationally efficient collection of translated and adapted QA samples across domains such as general knowledge, science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_pt.simple_bench
📊 Simple Bench Dataset
A Compact Benchmark for Structured Reasoning and Multiple-Choice Evaluation in Large Language Models
Simple Bench Dataset is a structured evaluation collection derived from the Simple Bench benchmark, designed to assess reasoning, comprehension, and multiple-choice question-answering capabilities of large language models through concise yet non-trivial problems that require logical inference rather than simple retrieval; each sample consists of a natural… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/simple_bench.indaqa
INDAQA - Italian Narrative Dataset for Long-document Question-Answering
INDAQA is the first Italian question-answering dataset specifically designed for long-context Italian narrative texts.
The dataset contains 362 documents paired with reading comprehension questions and reference answers based on Italian literary works sourced from Wikisource.
Questions and answers were automatically generated using Gemini and subsequently underwent both automatic filtering and… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/indaqa.enem_2025
🇧🇷 ENEM 2025 — Brazilian National High School Exam Dataset
A High-Quality Benchmark for Portuguese Academic Reasoning in Large Language Models
ENEM 2025 Dataset is a curated collection of question-answer pairs derived from the 2025 edition of the Brazilian National High School Exam (ENEM), designed to evaluate and improve the reasoning, reading comprehension, and multiple-choice answering capabilities of large language models in Brazilian Portuguese; as one of the largest… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/enem_2025.global_mmlu_lite
🌍 Global-MMLU Lite Dataset
A Lightweight Benchmark for Multi-Domain Reasoning in Large Language Models
Global-MMLU Lite is a curated and efficient subset of the Global Massive Multitask Language Understanding (MMLU) benchmark, designed to evaluate and fine-tune large language models across a wide range of academic and professional domains through high-quality multiple-choice question answering; preserving the diversity and rigor of the original benchmark while significantly… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite.global_mmlu_lite_en
🌍 Global-MMLU Lite (English Only)
A Focused Benchmark for English-Language Reasoning in Large Language Models
Global-MMLU Lite (English Only) is a curated subset of the Global-MMLU Lite benchmark specifically designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models within the English language, providing a diverse yet computationally efficient collection of structured QA samples spanning domains such as science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_en.global_mmlu_lite_es
🌎 Global-MMLU Lite (Spanish Only)
A Focused Benchmark for Spanish-Language Reasoning in Large Language Models
Global-MMLU Lite (Spanish Only) is a curated subset of the Global-MMLU Lite benchmark specifically designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models in Spanish, providing a diverse and computationally efficient collection of fully translated and standardized QA samples across domains such as science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_es.
