CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sapientinc /sudoku-extreme Hardest Sudoku Puzzle Dataset V2 This dataset contains a mixture of easy and very hard Sudoku puzzles collected from the Sudoku community. Dataset Composition Sources tdoku benchmarks enjoysudoku Easy Puzzles (1.1M) puzzles0_kaggle puzzles1_unbiased puzzles2_17_clue Hard Puzzles (3.1M) puzzles3_magictour_top1465 puzzles4_forum_hardest_1905 puzzles6_forum_hardest_1106 ph_2010/01_file1.txt Dataset Characteristics All… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/sudoku-extreme.textquestion-answering1M<n<10M35 likes4k downloads2y agoHugging Face02sapientinc /HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts. Citation If you find this project or our paper useful, please consider citing our paper: @misc{wang2026hrmtextefficientpretrainingscaling, title={HRM-Text: Efficient Pretraining Beyond Scaling}, author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.texttext-generation100M<n<1B17 likes3.9k downloads4mo agoHugging Face03sapinsapin /filipinospeechcorpus Filipino Speech Corpus (FSC) Studio-recorded Filipino read, spontaneous, and word-level speech — 125 speakers, packaged as ready-to-stream Parquet. 313,322 transcribed segments · 65.1 hours · 125 speakers · 16kHz mono This is the Filipino Speech Corpus (Sagum), recorded in a controlled setting and hand/machine transcribed with Transcriber. This repo repackages the original .wav + .trs volumes as segment-level Parquet with inline audio, so you can stream it without… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/filipinospeechcorpus.audioautomatic-speech-recognition100K<n<1M3 likes602 downloads1mo agoHugging Face04sapientinc /maze-30x30-hard-1ktabular1K<n<10K7 likes598 downloads1y agoHugging Face05sapienzanlp /dromedario-3-sft-dataset 🐪 Dataset Card for Dromedario 3 📋 Dataset Summary Dromedario 3 is a large-scale Italian instruction-tuning dataset derived from the English Tülu 3 SFT mixture through a principled translation pipeline: instructions and responses are classified according to the Natural Instructions taxonomy, classes are manually reviewed for translation safety, and items in validated classes are machine-translated into Italian. The full procedure is described in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/dromedario-3-sft-dataset.texttext-generation100K<n<1M9 likes456 downloads6d agoHugging Face06sapinsapin /pld Philippine Language Dataset (PLD) Ten Philippine languages, 980 speakers, 448 hours of prompted speech — one of the largest multilingual Philippine speech collections available as Parquet. 334,268 utterances · 448.2 hours · 980 speakers · 10 languages · 16kHz mono ▶ Try the models in your browser — transcribe, synthesize, or convert a voice in any of the ten languages, from your microphone or the preloaded clips. Collected by the University of the Philippines Diliman… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/pld.audioautomatic-speech-recognition100K<n<1M0 likes450 downloads1mo agoHugging Face07sapientinc /sudoku-extreme-1ktexttranslation10K<n<100K3 likes346 downloads1y agoHugging Face08sapienzanlp /wic Word in Context (WIC) Original Paper: https://wic-ita.github.io/ This dataset comes from EVALITA-2023. Word in Context task consists of establishing if a word w occurring in two different sentences s1 and s2 has the same meaning or not. We repropose this task to test generative LLMs defining a specific prompting strategy comparing the perplexities of possible continuations to understand the models' capabilities. Example Here you can see the structure of the single… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/wic.tabular1K<n<10K0 likes284 downloads2y agoHugging Face09schneiderkamplab /sapient-synth-tasksource-reclor sapient-synth-tasksource-reclor Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 4633 Task: synthetic anonymous instruction replacement Generation Rows were generated with google/gemma-4-31B-it and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-tasksource-reclor.text1K<n<10K0 likes277 downloads3mo agoHugging Face10viridono /CF-MS_Homo_sapiens_PPI CF-MS Elution Profile PPI Dataset Proteins typically function as part of larger complexes, and co-fractionation mass spectrometry (CF-MS) identifies these complexes by tracking which proteins "co-elute" — separate into the same fractions — during chromatography, since interacting proteins show highly correlated abundance patterns across fractions. These correlations are conventionally scored with a linear metric (Pearson correlation), but non-linear relationships in the elution… See the full description on the dataset page: https://huggingface.co/datasets/viridono/CF-MS_Homo_sapiens_PPI.text10M<n<100M2 likes274 downloads11d agoHugging Face11sapinsapin /halo-hil halo-hil Dataset Summary halo-hil is a web-scraped hil text corpus assembled for LLM pre-training. It contains documents from news sites, blogs, academic journals, and other web sources. Cleaning Pipeline The raw text column contains web-scraped content with significant noise. A cleaning pipeline produces the text_cleaned column by: Dropping navigation menus, markdown tables, bare URLs, image markdown Removing WordPress, Blogger, Scribd, and SlideShare… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halo-hil.text1K<n<10K0 likes274 downloads6mo agoHugging Face12songlab /gpn-msa-sapiens-dataset Training windows for GPN-MSA-Sapiens For more information check out our paper and repository. Path in Snakemake: results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001 tabular1M<n<10M0 likes242 downloads2y agoHugging Face13schneiderkamplab /sapient-synth-platypus-reclor sapient-synth-platypus-reclor Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 5131 Task: synthetic anonymous instruction replacement Generation Rows were generated with google/gemma-4-31B-it and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-platypus-reclor.text1K<n<10K0 likes167 downloads3mo agoHugging Face14sapienzanlp /mmlu_italian MMLU - Italian (IT) This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics. Dataset Details The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/mmlu_italian.texttext-generation10K<n<100K1 likes141 downloads10mo agoHugging Face15sapienzanlp /boolq_italian BoolQ - Italian (IT) This dataset is an Italian translation of BoolQ. BoolQ is a question-answering dataset composed of user queries issued to a search engine. Dataset Details The task is to predict whether the answer to the question is true or false based on the context provided in the question. A text snippet from Wikipedia is provided as the context for each question. The dataset includes the following splits: Train: 9,427 rows Validation: 3,270 rows… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/boolq_italian.texttext-generation10K<n<100K0 likes109 downloads10mo agoHugging Face16sapienzanlp /arc_italian ARC - Italian (IT) This dataset is an Italian translation of the AI2 Reasoning Challenge (ARC). ARC is a question-answering dataset that requires an understanding of natural language text and reasoning capabilities to answer questions correctly. Dataset Details The dataset consists of multiple-choice questions, where each question is associated with a set of answer choices (up to 5 choices). The task is to choose the correct answer choice based on the context provided in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/arc_italian.texttext-generation1K<n<10K2 likes105 downloads10mo agoHugging Face17sapienzanlp /zebra-kb-explanations ZEBRA: Zero-Shot Example-Based Retrieval Augmentation for Commonsense Question Answering                     A retrieval augmentation framework for zero-shot commonsense question answering with LLMs. 🛠️ Installation Installation from PyPi pip install zebra-qa Installation from source git clone https://github.com/sapienzanlp/zebra.git cd zebra conda create -n zebra python==3.10 conda activate zebra pip install -e . 🚀 Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/zebra-kb-explanations.text100K<n<1M3 likes97 downloads2y agoHugging Face18sapienzanlp /gsm8k_italian GSM8K - Italian (IT) This dataset is an Italian translation of GSM8K. GSM8K stands for Grade School Math 8K, a dataset for math word problems, which should be easy to solve for people with an elementary school education. Dataset Details The dataset consists of math word problems, where each problem is associated with a possible explanation of how to solve it. The task is to generate the answer to the math problem. The dataset is split into a training set and a test set.… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/gsm8k_italian.texttext-generation1K<n<10K1 likes96 downloads10mo agoHugging Face19sapienzanlp /hellaswag_italian HellaSwag - Italian (IT) This dataset is an Italian translation of HellaSwag. HellaSwag is a large-scale commonsense reasoning dataset, which requires reading comprehension and commonsense reasoning to predict the correct ending of a sentence. Dataset Details The dataset consists of instances containing a context and a multiple-choice question with four possible answers. The task is to predict the correct ending of the sentence. The dataset is split into a training set… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/hellaswag_italian.texttext-generation10K<n<100K1 likes95 downloads10mo agoHugging Face20sapienzanlp /ea-mt-benchmark Dataset Card for EA-MT EA-MT (Entity-Aware Machine Translation) is a multilingual benchmark for evaluating the capabilities of Large Language Models (LLMs) and Machine Translation (MT) models in translating simple sentences with potentially challenging entity mentions, e.g., entities for which a word-for-word translation may not be accurate. Here is an example of a simple sentence with a challenging entity mention: English: "What is the plot of The Catcher in the Rye?" Italian:… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ea-mt-benchmark.texttext-generation10K<n<100K6 likes83 downloads2y agoHugging Face21sapienzanlp /INDAQA_CALAMITA Dataset Card for INDAQA 2 INDAQA 2 (CALAMITA update) is a large-scale Italian reading-comprehension and question-answering benchmark built from classic narrative works. The dataset is designed to support research in Italian NLP, reading comprehension, information retrieval, and language model evaluation on medium- and long-context narratives and it is released as part of the CALAMITA 2026 edition. Dataset Details Dataset Description INDAQA 2 is a… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/INDAQA_CALAMITA.textn<1K2 likes81 downloads7mo agoHugging Face22sapienzanlp-course-materials /hw-mnlp-2026 Dataset for Multilingual Natural Language Processing (MNLP) Homeworks This dataset serves for both Homework 1 and Homework 2 of the Multilingual Natural Language Processing (MNLP) course. Homework 1 - Semantic Search In the first homework, you are asked to build semantic search systems. You must only use the following variables: query: A single question in natural language. query_id: The question (query) identifier. candidate_chunks: List of candidate answers (only one… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp-course-materials/hw-mnlp-2026.tabularsentence-similarity10K<n<100K0 likes78 downloads6mo agoHugging Face23sapienzanlp /ReTraceQA Dataset Card for ReTraceQA Dataset Summary ReTraceQA is a dataset designed to evaluate the reasoning traces of Small Language Models (SLMs) on commonsense reasoning tasks. It includes model-generated traces across four benchmark datasets: CommonsenseQA, OpenBookQA, QASC, and StrategyQA. During the construction of ReTraceQA, only correct instances from the original benchmarks were retained, and erroneous instances were manually removed to ensure data quality. Each item in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ReTraceQA.tabular1K<n<10K1 likes73 downloads5mo agoHugging Face24sapienzanlp /BBH_italian BBH - Italian (IT) This dataset is an Italian translation of the BBH dataset. BBH Bench dataset and consists of 23 tasks that are particularly hard for current generation of language models. The dataset is called Big Bench Hard. Boolean Expressions: Evaluate the truth value of a random Boolean expression consisting of Boolean constants (True, False) and basic Boolean operators (and, or, not). Causal Judgment: Given a short story (involving moral, intentional, or counterfactual… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/BBH_italian.text1K<n<10K0 likes71 downloads10mo agoHugging Face25sapienzanlp /prelearn Prerequisite RElation LEARNing (PRELEARN) Original Paper: https://ceur-ws.org/Vol-2765/paper164.pdf This dataset contains a collection of binary-labelled concept pairs (A,B) extracted from textbooks on four domains: data mining, geometry, physics and precalculus. Then, domain experts were asked to manually annotate if pairs of concepts showed a prerequisite relation or not, therefore the dataset consists of both positive and negative concept pairs. We obtained the data from the… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/prelearn.text1K<n<10K1 likes69 downloads2y agoHugging Face26sapienzanlp /ITALIC-gen Dataset Card for ITALICGEN ITALICGEN is an adaptation of ITALIC (a Multiple-choice QA (MCQA) benchmark focused on the Italian culture) to a generative, Open-ended (OE) setting.Note: The sample in the figure is a direct translation; the original questions are in Italian. Dataset Details Dataset Description ITALICGEN is entirely based on ITALIC; for in-depth details, refer to the original publication (Seveso et al., 2025, ITALIC: An Italian… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ITALIC-gen.textquestion-answering1K<n<10K5 likes68 downloads10mo agoHugging Face27sapienzanlp /quandho QUANDHO: QUestion ANswering Data for italian HistOry Original Paper: https://aclanthology.org/L16-1069.pdf QUANDHO (QUestion ANswering Data for italian HistOry) is an Italian question answering dataset created to cover the history of Italy in the first half of the XX century. Starting from QUANDHO we defined a Multi-choice QA dataset, with a correct answer and four different distractors. Data and Distractors Generation We relied on the original data, to create this… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/quandho.text1K<n<10K1 likes60 downloads2y agoHugging Face28sapienzanlp /piqa_italian PIQA - Italian (IT) This dataset is an Italian translation of PIQA. PIQA stands for Physical Interaction Question Answering, a dataset of questions about common scenarios that require an understanding of the physical world. Dataset Details The dataset consists of questions about common scenarios that require an understanding of the physical world. Each question is associated with a correct answer and a distractor. The task is to predict the correct answer to the… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/piqa_italian.texttext-generation10K<n<100K0 likes57 downloads10mo agoHugging Face29sapienzanlp /MMLU-Adversarial Dataset Card for MMLU-Adversarial Dataset Summary MMLU-Adversarial is a diagnostic dataset designed to evaluate the ability of current LLM-based answer extraction techniques to detect instances in which the model produces invalid answers due to hallucinated or flawed reasoning. Each instance in the dataset includes a reasoning chain that undermines the validity of the final selected answer, and as such, should be labeled as invalid (e.g., [No Valid Answer]). The flawed… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/MMLU-Adversarial.text1K<n<10K1 likes56 downloads1y agoHugging Face30sapienzanlp /sciq_italian SciQ - Italian (IT) This dataset is an Italian translation of SciQ. SciQ is a dataset for scientific questions, which were semi-automatically generated from an existing set of questions. The dataset is designed to test the ability of models to answer questions that require scientific knowledge. Dataset Details The dataset consists of science-related questions, where each question is associated with a correct answer and three possible distractors. The task is to predict… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/sciq_italian.texttext-generation1K<n<10K0 likes52 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.