CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01schneiderkamplab /sapient-synth-tasksource-reclor sapient-synth-tasksource-reclor Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 4633 Task: synthetic anonymous instruction replacement Generation Rows were generated with google/gemma-4-31B-it and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-tasksource-reclor.text1K<n<10K0 likes277 downloads3mo agoHugging Face02schneiderkamplab /sapient-synth-platypus-reclor sapient-synth-platypus-reclor Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 5131 Task: synthetic anonymous instruction replacement Generation Rows were generated with google/gemma-4-31B-it and… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-platypus-reclor.text1K<n<10K0 likes167 downloads3mo agoHugging Face03sapienzanlp /mmlu_italian MMLU - Italian (IT) This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics. Dataset Details The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/mmlu_italian.texttext-generation10K<n<100K1 likes141 downloads10mo agoHugging Face04sapienzanlp /boolq_italian BoolQ - Italian (IT) This dataset is an Italian translation of BoolQ. BoolQ is a question-answering dataset composed of user queries issued to a search engine. Dataset Details The task is to predict whether the answer to the question is true or false based on the context provided in the question. A text snippet from Wikipedia is provided as the context for each question. The dataset includes the following splits: Train: 9,427 rows Validation: 3,270 rows… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/boolq_italian.texttext-generation10K<n<100K0 likes109 downloads10mo agoHugging Face05sapienzanlp /arc_italian ARC - Italian (IT) This dataset is an Italian translation of the AI2 Reasoning Challenge (ARC). ARC is a question-answering dataset that requires an understanding of natural language text and reasoning capabilities to answer questions correctly. Dataset Details The dataset consists of multiple-choice questions, where each question is associated with a set of answer choices (up to 5 choices). The task is to choose the correct answer choice based on the context provided in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/arc_italian.texttext-generation1K<n<10K2 likes105 downloads10mo agoHugging Face06sapienzanlp /gsm8k_italian GSM8K - Italian (IT) This dataset is an Italian translation of GSM8K. GSM8K stands for Grade School Math 8K, a dataset for math word problems, which should be easy to solve for people with an elementary school education. Dataset Details The dataset consists of math word problems, where each problem is associated with a possible explanation of how to solve it. The task is to generate the answer to the math problem. The dataset is split into a training set and a test set.… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/gsm8k_italian.texttext-generation1K<n<10K1 likes96 downloads10mo agoHugging Face07sapienzanlp /hellaswag_italian HellaSwag - Italian (IT) This dataset is an Italian translation of HellaSwag. HellaSwag is a large-scale commonsense reasoning dataset, which requires reading comprehension and commonsense reasoning to predict the correct ending of a sentence. Dataset Details The dataset consists of instances containing a context and a multiple-choice question with four possible answers. The task is to predict the correct ending of the sentence. The dataset is split into a training set… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/hellaswag_italian.texttext-generation10K<n<100K1 likes95 downloads10mo agoHugging Face08sapienzanlp /ea-mt-benchmark Dataset Card for EA-MT EA-MT (Entity-Aware Machine Translation) is a multilingual benchmark for evaluating the capabilities of Large Language Models (LLMs) and Machine Translation (MT) models in translating simple sentences with potentially challenging entity mentions, e.g., entities for which a word-for-word translation may not be accurate. Here is an example of a simple sentence with a challenging entity mention: English: "What is the plot of The Catcher in the Rye?" Italian:… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ea-mt-benchmark.texttext-generation10K<n<100K6 likes83 downloads2y agoHugging Face09sapienzanlp /piqa_italian PIQA - Italian (IT) This dataset is an Italian translation of PIQA. PIQA stands for Physical Interaction Question Answering, a dataset of questions about common scenarios that require an understanding of the physical world. Dataset Details The dataset consists of questions about common scenarios that require an understanding of the physical world. Each question is associated with a correct answer and a distractor. The task is to predict the correct answer to the… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/piqa_italian.texttext-generation10K<n<100K0 likes57 downloads10mo agoHugging Face10sapienzanlp /sciq_italian SciQ - Italian (IT) This dataset is an Italian translation of SciQ. SciQ is a dataset for scientific questions, which were semi-automatically generated from an existing set of questions. The dataset is designed to test the ability of models to answer questions that require scientific knowledge. Dataset Details The dataset consists of science-related questions, where each question is associated with a correct answer and three possible distractors. The task is to predict… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/sciq_italian.texttext-generation1K<n<10K0 likes52 downloads10mo agoHugging Face11sapienzanlp /winogrande_italian Winogrande - Italian (IT) This dataset is an Italian translation of Winogrande. Winogrande is a large-scale dataset for coreference resolution, commonsense reasoning, and world knowledge. It is based on the original Winograd Schema Challenge dataset. Dataset Details The dataset consists of almost 40K examples, each containing a sentence with a blank and two possible fill-in-the-blank options. The task is to choose the correct option that correctly fills in the blank based… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/winogrande_italian.texttext-generation1K<n<10K0 likes50 downloads10mo agoHugging Face12sapienzanlp /truthful_qa_italian TruthfulQA - Italian (IT) This dataset is an Italian translation of TruthfulQA. TruthfulQA is a dataset for fact-based question answering, which contains questions that require factual knowledge to answer correctly. These questions are designed so that some humans would answer them incorrectly because of common misconceptions. Dataset Details The dataset is a question answering dataset that contains questions that require factual knowledge to answer correctly and avoid… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/truthful_qa_italian.texttext-generationn<1K1 likes50 downloads10mo agoHugging Face13brozonoyer /sapientinc-sudoku-extreme-timvink-sudoku-solver Dataset Card for Dataset Name This dataset is built on top of sapientinc/sudoku-extreme by preprocessing every example with timvink/sudoku-solver. The additional fields are: trajectory: the step-by-step sudoku board configuration from the question to the solution, acquired by running puzzle.solve_step() until complete num_steps: number of calls to puzzle.solve_step() required to obtain solution strategies_used: set of strategies used by the solver tabular1M<n<10M0 likes49 downloads8mo agoHugging Face14sapienzanlp /nounatlas_srl_corpus NounAtlas SRL Corpus This dataset is part of the NounAtlas project, aiming to enhance Nominal Semantic Role Labeling (SRL) by providing a comprehensive inventory of nominal predicates organized into semantically-coherent frames. Dataset Details The NounAtlas SRL Corpus contains sentences annotated with nominal predicates and their corresponding semantic roles. This dataset is split into three subsets: training, development, and test. Train: 22,452 sentences Dev: 2,806… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/nounatlas_srl_corpus.texttoken-classification10K<n<100K1 likes47 downloads2y agoHugging Face15TooKeen /sapientblock-blockchain-use-cases SapientBlock Blockchain Use Cases Der Datensatz enthält 255 redaktionell geprüfte Blockchain-Use-Cases aus 74 Branchen. Er stellt die öffentlich zugänglichen SapientBlock-Inhalte in einem maschinenlesbaren JSONL-Format für Forschung, Bildung, Retrieval und Quellenanalyse bereit. SapientBlock ist ein öffentliches Forschungs- und Bildungsprojekt von ShapeNeural. Die Inhalte sind keine Rechts-, Investitions-, Unternehmens- oder technische Beratung. Inhalt Jeder… See the full description on the dataset page: https://huggingface.co/datasets/TooKeen/sapientblock-blockchain-use-cases.textn<1K0 likes47 downloads4d agoHugging Face16sapiens-technology /global_mmlu_lite_pt 🌎 Global-MMLU Lite (Portuguese) A Focused Benchmark for Portuguese-Language Reasoning in Large Language Models Global-MMLU Lite (Portuguese) is a curated subset of the Global-MMLU Lite benchmark designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models in Portuguese, providing a diverse and computationally efficient collection of translated and adapted QA samples across domains such as general knowledge, science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_pt.textquestion-answeringn<1K0 likes46 downloads5mo agoHugging Face17sapiens-technology /enem_2025 🇧🇷 ENEM 2025 — Brazilian National High School Exam Dataset A High-Quality Benchmark for Portuguese Academic Reasoning in Large Language Models ENEM 2025 Dataset is a curated collection of question-answer pairs derived from the 2025 edition of the Brazilian National High School Exam (ENEM), designed to evaluate and improve the reasoning, reading comprehension, and multiple-choice answering capabilities of large language models in Brazilian Portuguese; as one of the largest… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/enem_2025.textquestion-answering1K<n<10K0 likes35 downloads5mo agoHugging Face18sapiens-technology /simple_bench 📊 Simple Bench Dataset A Compact Benchmark for Structured Reasoning and Multiple-Choice Evaluation in Large Language Models Simple Bench Dataset is a structured evaluation collection derived from the Simple Bench benchmark, designed to assess reasoning, comprehension, and multiple-choice question-answering capabilities of large language models through concise yet non-trivial problems that require logical inference rather than simple retrieval; each sample consists of a natural… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/simple_bench.texttext-generationn<1K0 likes35 downloads5mo agoHugging Face19schneiderkamplab /sapient-synth-flan-niv2-fsopt-data-task902-deceptive-opinion-spam-classification sapient-synth-flan-niv2-fsopt-data-task902-deceptive-opinion-spam-classification Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 2935 Task: synthetic anonymous instruction replacement Generation… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task902-deceptive-opinion-spam-classification.text1K<n<10K0 likes33 downloads3mo agoHugging Face20schneiderkamplab /sapient-synth-flan-niv2-fsopt-data-task635-allegro-reviews-answer-generation sapient-synth-flan-niv2-fsopt-data-task635-allegro-reviews-answer-generation Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 3186 Task: synthetic anonymous instruction replacement Generation Rows… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task635-allegro-reviews-answer-generation.text1K<n<10K0 likes28 downloads3mo agoHugging Face21sapiens-technology /global_mmlu_lite 🌍 Global-MMLU Lite Dataset A Lightweight Benchmark for Multi-Domain Reasoning in Large Language Models Global-MMLU Lite is a curated and efficient subset of the Global Massive Multitask Language Understanding (MMLU) benchmark, designed to evaluate and fine-tune large language models across a wide range of academic and professional domains through high-quality multiple-choice question answering; preserving the diversity and rigor of the original benchmark while significantly… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite.textquestion-answering10K<n<100K0 likes25 downloads5mo agoHugging Face22schneiderkamplab /sapient-synth-flan-niv2-fsopt-data-task1377-newscomm-translation sapient-synth-flan-niv2-fsopt-data-task1377-newscomm-translation Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 3491 Task: synthetic anonymous instruction replacement Generation Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task1377-newscomm-translation.text1K<n<10K0 likes25 downloads3mo agoHugging Face23sapiens-technology /global_mmlu_lite_en 🌍 Global-MMLU Lite (English Only) A Focused Benchmark for English-Language Reasoning in Large Language Models Global-MMLU Lite (English Only) is a curated subset of the Global-MMLU Lite benchmark specifically designed to evaluate the reasoning, knowledge, and multiple-choice question-answering capabilities of large language models within the English language, providing a diverse yet computationally efficient collection of structured QA samples spanning domains such as science… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/global_mmlu_lite_en.textquestion-answeringn<1K0 likes18 downloads5mo agoHugging Face24schneiderkamplab /sapient-synth-tasksource-pragmeval-sarcasm sapient-synth-tasksource-pragmeval-sarcasm Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 3753 Task: synthetic anonymous instruction replacement Generation Rows were generated with… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-tasksource-pragmeval-sarcasm.text1K<n<10K0 likes17 downloads3mo agoHugging Face25schneiderkamplab /sapient-synth-flan-niv2-fsopt-data-task903-deceptive-opinion-spam-classification sapient-synth-flan-niv2-fsopt-data-task903-deceptive-opinion-spam-classification Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 1703 Task: synthetic anonymous instruction replacement Generation… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-fsopt-data-task903-deceptive-opinion-spam-classification.text1K<n<10K0 likes16 downloads3mo agoHugging Face26sapiens-technology /math_precision_benchmarking 🧮 Math Precision — Benchmarking A Formal Framework for High-Precision Arithmetic Evaluation in Large Language Models Math Precision — Benchmarking, developed by Sapiens Technology®, is a rigorous framework for evaluating the true arithmetic capabilities of large language models by generating fully stochastic, high-precision mathematical problems that eliminate memorization and heuristic guessing; operating in a 100-digit floating-point field, it forces extreme numerical precision… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/math_precision_benchmarking.textn<1K0 likes15 downloads5mo agoHugging Face27schneiderkamplab /sapient-synth-flan-dialog-fsopt-data-qrecc sapient-synth-flan-dialog-fsopt-data-qrecc Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 20748 Task: synthetic anonymous instruction replacement Generation Rows were generated with… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-dialog-fsopt-data-qrecc.text10K<n<100K0 likes15 downloads3mo agoHugging Face28schneiderkamplab /sapient-synth-flan-flan-zsopt-data-opinion-abstracts-idebate sapient-synth-flan-flan-zsopt-data-opinion-abstracts-idebate Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 1307 Task: synthetic anonymous instruction replacement Generation Rows were generated… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-flan-zsopt-data-opinion-abstracts-idebate.text1K<n<10K0 likes15 downloads3mo agoHugging Face29schneiderkamplab /sapient-synth-flan-niv2-zsopt-data-task1375-newscomm-translation sapient-synth-flan-niv2-zsopt-data-task1375-newscomm-translation Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 2180 Task: synthetic anonymous instruction replacement Generation Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-zsopt-data-task1375-newscomm-translation.text1K<n<10K0 likes15 downloads3mo agoHugging Face30schneiderkamplab /sapient-synth-flan-niv2-zsopt-data-task1371-newscomm-translation sapient-synth-flan-niv2-zsopt-data-task1371-newscomm-translation Chat-template-ready synthetic anonymous replacement examples for one Sapient source excluded from the DFM5 data mix. Contents Format: gzip-compressed JSON Lines under data/train.jsonl.gz Schema: {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]} Files: 1 Rows: 1577 Task: synthetic anonymous instruction replacement Generation Rows were… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/sapient-synth-flan-niv2-zsopt-data-task1371-newscomm-translation.text1K<n<10K0 likes14 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.