CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sapientinc /sudoku-extreme Hardest Sudoku Puzzle Dataset V2 This dataset contains a mixture of easy and very hard Sudoku puzzles collected from the Sudoku community. Dataset Composition Sources tdoku benchmarks enjoysudoku Easy Puzzles (1.1M) puzzles0_kaggle puzzles1_unbiased puzzles2_17_clue Hard Puzzles (3.1M) puzzles3_magictour_top1465 puzzles4_forum_hardest_1905 puzzles6_forum_hardest_1106 ph_2010/01_file1.txt Dataset Characteristics All… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/sudoku-extreme.textquestion-answering1M<n<10M35 likes4k downloads2y agoHugging Face02sapienzanlp-course-materials /hw-mnlp-2026 Dataset for Multilingual Natural Language Processing (MNLP) Homeworks This dataset serves for both Homework 1 and Homework 2 of the Multilingual Natural Language Processing (MNLP) course. Homework 1 - Semantic Search In the first homework, you are asked to build semantic search systems. You must only use the following variables: query: A single question in natural language. query_id: The question (query) identifier. candidate_chunks: List of candidate answers (only one… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp-course-materials/hw-mnlp-2026.tabularsentence-similarity10K<n<100K0 likes78 downloads6mo agoHugging Face03sapienzanlp /ITALIC-gen Dataset Card for ITALICGEN ITALICGEN is an adaptation of ITALIC (a Multiple-choice QA (MCQA) benchmark focused on the Italian culture) to a generative, Open-ended (OE) setting.Note: The sample in the figure is a direct translation; the original questions are in Italian. Dataset Details Dataset Description ITALICGEN is entirely based on ITALIC; for in-depth details, refer to the original publication (Seveso et al., 2025, ITALIC: An Italian… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ITALIC-gen.textquestion-answering1K<n<10K5 likes64 downloads10mo agoHugging Face04sapiens-technology /global_mmlu_lite_pt 🌎 Global-MMLU Lite (Portuguese) A Focused Benchmark for Portuguese-Language Reasoning in Large Language Models Global-MMLU Lite (Portuguese) is a curated subset of the Global-MMLU Lite benchmark designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models in Portuguese, providing a diverse and computationally efficient collection of translated and adapted QA samples across domains such as general knowledge, science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_pt.textquestion-answeringn<1K0 likes46 downloads5mo agoHugging Face05sapiens-technology /simple_bench 📊 Simple Bench Dataset A Compact Benchmark for Structured Reasoning and Multiple-Choice Evaluation in Large Language Models Simple Bench Dataset is a structured evaluation collection derived from the Simple Bench benchmark, designed to assess reasoning, comprehension, and multiple-choice question-answering capabilities of large language models through concise yet non-trivial problems that require logical inference rather than simple retrieval; each sample consists of a natural… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/simple_bench.texttext-generationn<1K0 likes38 downloads5mo agoHugging Face06sapienzanlp /indaqa INDAQA - Italian Narrative Dataset for Long-document Question-Answering INDAQA is the first Italian question-answering dataset specifically designed for long-context Italian narrative texts. The dataset contains 362 documents paired with reading comprehension questions and reference answers based on Italian literary works sourced from Wikisource. Questions and answers were automatically generated using Gemini and subsequently underwent both automatic filtering and… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/indaqa.tabularquestion-answeringn<1K3 likes32 downloads8mo agoHugging Face07sapiens-technology /enem_2025 🇧🇷 ENEM 2025 — Brazilian National High School Exam Dataset A High-Quality Benchmark for Portuguese Academic Reasoning in Large Language Models ENEM 2025 Dataset is a curated collection of question-answer pairs derived from the 2025 edition of the Brazilian National High School Exam (ENEM), designed to evaluate and improve the reasoning, reading comprehension, and multiple-choice answering capabilities of large language models in Brazilian Portuguese; as one of the largest… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/enem_2025.textquestion-answering1K<n<10K0 likes27 downloads5mo agoHugging Face08sapiens-technology /global_mmlu_lite 🌍 Global-MMLU Lite Dataset A Lightweight Benchmark for Multi-Domain Reasoning in Large Language Models Global-MMLU Lite is a curated and efficient subset of the Global Massive Multitask Language Understanding (MMLU) benchmark, designed to evaluate and fine-tune large language models across a wide range of academic and professional domains through high-quality multiple-choice question answering; preserving the diversity and rigor of the original benchmark while significantly… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite.textquestion-answering10K<n<100K0 likes18 downloads5mo agoHugging Face09sapiens-technology /global_mmlu_lite_en 🌍 Global-MMLU Lite (English Only) A Focused Benchmark for English-Language Reasoning in Large Language Models Global-MMLU Lite (English Only) is a curated subset of the Global-MMLU Lite benchmark specifically designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models within the English language, providing a diverse yet computationally efficient collection of structured QA samples spanning domains such as science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_en.textquestion-answeringn<1K0 likes18 downloads5mo agoHugging Face10sapiens-technology /global_mmlu_lite_es 🌎 Global-MMLU Lite (Spanish Only) A Focused Benchmark for Spanish-Language Reasoning in Large Language Models Global-MMLU Lite (Spanish Only) is a curated subset of the Global-MMLU Lite benchmark specifically designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models in Spanish, providing a diverse and computationally efficient collection of fully translated and standardized QA samples across domains such as science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_es.textquestion-answeringn<1K0 likes10 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.