CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sapientinc /HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts. Citation If you find this project or our paper useful, please consider citing our paper: @misc{wang2026hrmtextefficientpretrainingscaling, title={HRM-Text: Efficient Pretraining Beyond Scaling}, author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.texttext-generation100M<n<1B17 likes3.5k downloads4mo agoHugging Face02sapienzanlp /dromedario-3-sft-dataset 🐪 Dataset Card for Dromedario 3 📋 Dataset Summary Dromedario 3 is a large-scale Italian instruction-tuning dataset derived from the English Tülu 3 SFT mixture through a principled translation pipeline: instructions and responses are classified according to the Natural Instructions taxonomy, classes are manually reviewed for translation safety, and items in validated classes are machine-translated into Italian. The full procedure is described in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/dromedario-3-sft-dataset.texttext-generation100K<n<1M9 likes502 downloads8d agoHugging Face03sapinsapin /halo-hil halo-hil Web text in hil, re-filtered by language and prepared for pretraining. What changed, and why it had to The earlier version of this dataset was labelled hil by the crawler's own language detection, and that label was never verified. An audit on 2026-09-22 found that most of it was not hil: over a random sample of 1,499 sentences, GlotLID v3 called 44 % English, 22 % Filipino/Tagalog and only 12 % Hiligaynon — much of the corpus was Tagalog news copy and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halo-hil.tabulartext-generationn<1K0 likes357 downloads1d agoHugging Face04sapienzanlp /mmlu_italian MMLU - Italian (IT) This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics. Dataset Details The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/mmlu_italian.texttext-generation10K<n<100K1 likes165 downloads10mo agoHugging Face05sapienzanlp /arc_italian ARC - Italian (IT) This dataset is an Italian translation of the AI2 Reasoning Challenge (ARC). ARC is a question-answering dataset that requires an understanding of natural language text and reasoning capabilities to answer questions correctly. Dataset Details The dataset consists of multiple-choice questions, where each question is associated with a set of answer choices (up to 5 choices). The task is to choose the correct answer choice based on the context provided in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/arc_italian.texttext-generation1K<n<10K2 likes120 downloads10mo agoHugging Face06sapienzanlp /boolq_italian BoolQ - Italian (IT) This dataset is an Italian translation of BoolQ. BoolQ is a question-answering dataset composed of user queries issued to a search engine. Dataset Details The task is to predict whether the answer to the question is true or false based on the context provided in the question. A text snippet from Wikipedia is provided as the context for each question. The dataset includes the following splits: Train: 9,427 rows Validation: 3,270 rows… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/boolq_italian.texttext-generation10K<n<100K0 likes117 downloads10mo agoHugging Face07sapienzanlp /gsm8k_italian GSM8K - Italian (IT) This dataset is an Italian translation of GSM8K. GSM8K stands for Grade School Math 8K, a dataset for math word problems, which should be easy to solve for people with an elementary school education. Dataset Details The dataset consists of math word problems, where each problem is associated with a possible explanation of how to solve it. The task is to generate the answer to the math problem. The dataset is split into a training set and a test set.… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/gsm8k_italian.texttext-generation1K<n<10K1 likes104 downloads10mo agoHugging Face08sapienzanlp /ea-mt-benchmark Dataset Card for EA-MT EA-MT (Entity-Aware Machine Translation) is a multilingual benchmark for evaluating the capabilities of Large Language Models (LLMs) and Machine Translation (MT) models in translating simple sentences with potentially challenging entity mentions, e.g., entities for which a word-for-word translation may not be accurate. Here is an example of a simple sentence with a challenging entity mention: English: "What is the plot of The Catcher in the Rye?" Italian:… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ea-mt-benchmark.texttext-generation10K<n<100K6 likes98 downloads2y agoHugging Face09opendatalab /SA-Prot-annot SA-Prot-Annot Dataset (Sci-Align) 🌌 The Sciverse Data Foundation Sciverse is a comprehensive, multi-layered scientific data foundation designed to provide the ultimate data infrastructure for the AI for Science (AI4S) community. As scientific research becomes increasingly data-driven, Sciverse supplies the essential, high-quality data resources required to build robust scientific knowledge systems and accelerate research. Sciverse consists of three core data… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SA-Prot-annot.texttext-generation1M<n<10M4 likes98 downloads4mo agoHugging Face10Zichen1024 /SAP-9k SAP: State-Guided Data Synthesis with Argument Provenance for Multi-Turn Tool Use paper: SAP: State-Guided Data Synthesis with Argument Provenance for Multi-Turn Tool Use Dataset Overview This is a high-quality, executor-validated multi-turn tool-use dataset designed for training agentic language models on long-horizon function calling tasks. The dataset focuses on argument-level cross-turn dependency grounding, ensuring tool arguments are sourced from verifiable… See the full description on the dataset page: https://huggingface.co/datasets/Zichen1024/SAP-9k.texttext-generation1K<n<10K2 likes96 downloads11d agoHugging Face11sapienzanlp /hellaswag_italian HellaSwag - Italian (IT) This dataset is an Italian translation of HellaSwag. HellaSwag is a large-scale commonsense reasoning dataset, which requires reading comprehension and commonsense reasoning to predict the correct ending of a sentence. Dataset Details The dataset consists of instances containing a context and a multiple-choice question with four possible answers. The task is to predict the correct ending of the sentence. The dataset is split into a training set… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/hellaswag_italian.texttext-generation10K<n<100K1 likes94 downloads10mo agoHugging Face12sapienzanlp-course-materials /hw-mnlp-2026 Dataset for Multilingual Natural Language Processing (MNLP) Homeworks This dataset serves for both Homework 1 and Homework 2 of the Multilingual Natural Language Processing (MNLP) course. Homework 1 - Semantic Search In the first homework, you are asked to build semantic search systems. You must only use the following variables: query: A single question in natural language. query_id: The question (query) identifier. candidate_chunks: List of candidate answers (only one… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp-course-materials/hw-mnlp-2026.tabularsentence-similarity10K<n<100K0 likes71 downloads6mo agoHugging Face13sapbot /big-pickle-409xTrace of Big Pickle, stealth model from OpenCode Zen. (It's been rumored that it's GLM 4.6) Data is presented in ShareGPT format and each conversation split by newline. Ready to be used for fine-tuning. Brought to you by sapbot from Romarchive texttext-generationn<1K1 likes63 downloads5mo agoHugging Face14xiuwenz2 /SAP-Hypo5 Dataset Card for SAP-Hypo5 SAP-Hypo5 is an open benchmark for LLM-Agent ASR hypothesis correction on dysarthric speech. Dataset Description Following HyPoradise, each selected SAP utterance is paired with its reference transcript and the top-5 ASR hypotheses from Whisper-large-v2 fine-tuned on SAP (PD-only challenge release). Dataset Sources Repository: https://github.com/xiuwenz2/SAP-Hypo5 Paper: Towards Robust Dysarthric Speech Recognition: LLM-Agent… See the full description on the dataset page: https://huggingface.co/datasets/xiuwenz2/SAP-Hypo5.texttext-generation10K<n<100K0 likes57 downloads1y agoHugging Face15sapienzanlp /piqa_italian PIQA - Italian (IT) This dataset is an Italian translation of PIQA. PIQA stands for Physical Interaction Question Answering, a dataset of questions about common scenarios that require an understanding of the physical world. Dataset Details The dataset consists of questions about common scenarios that require an understanding of the physical world. Each question is associated with a correct answer and a distractor. The task is to predict the correct answer to the… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/piqa_italian.texttext-generation10K<n<100K0 likes51 downloads10mo agoHugging Face16sapienzanlp /truthful_qa_italian TruthfulQA - Italian (IT) This dataset is an Italian translation of TruthfulQA. TruthfulQA is a dataset for fact-based question answering, which contains questions that require factual knowledge to answer correctly. These questions are designed so that some humans would answer them incorrectly because of common misconceptions. Dataset Details The dataset is a question answering dataset that contains questions that require factual knowledge to answer correctly and avoid… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/truthful_qa_italian.texttext-generationn<1K1 likes50 downloads10mo agoHugging Face17sapienzanlp /winogrande_italian Winogrande - Italian (IT) This dataset is an Italian translation of Winogrande. Winogrande is a large-scale dataset for coreference resolution, commonsense reasoning, and world knowledge. It is based on the original Winograd Schema Challenge dataset. Dataset Details The dataset consists of almost 40K examples, each containing a sentence with a blank and two possible fill-in-the-blank options. The task is to choose the correct option that correctly fills in the blank based… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/winogrande_italian.texttext-generation1K<n<10K0 likes49 downloads10mo agoHugging Face18sapbot /yandexq-qa-chatmlChatML formatted version of its5Q/yandex-q. texttext-generation100K<n<1M1 likes48 downloads4mo agoHugging Face19sapienzanlp /sciq_italian SciQ - Italian (IT) This dataset is an Italian translation of SciQ. SciQ is a dataset for scientific questions, which were semi-automatically generated from an existing set of questions. The dataset is designed to test the ability of models to answer questions that require scientific knowledge. Dataset Details The dataset consists of science-related questions, where each question is associated with a correct answer and three possible distractors. The task is to predict… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/sciq_italian.texttext-generation1K<n<10K0 likes46 downloads10mo agoHugging Face20sapinsapin /BantayWika BantayWika A FineWeb-compatible pretraining text corpus for Philippine languages, derived from the Bantay-Wika corpus collected by the University of the Philippines Sentro ng Wikang Filipino (UP-SWF) and the UP Digital Signal Processing (DSP) Laboratory. The Bantay-Wika (Language Watch) project was started in 1994 by UP-SWF to track how the Philippine national language is used and develops, particularly in Philippine media. The first phase (1994–2004) involved manual collection and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/BantayWika.texttext-generation10K<n<100K0 likes46 downloads7mo agoHugging Face21sapiens-technology /simple_bench 📊 Simple Bench Dataset A Compact Benchmark for Structured Reasoning and Multiple-Choice Evaluation in Large Language Models Simple Bench Dataset is a structured evaluation collection derived from the Simple Bench benchmark, designed to assess reasoning, comprehension, and multiple-choice question-answering capabilities of large language models through concise yet non-trivial problems that require logical inference rather than simple retrieval; each sample consists of a natural… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/simple_bench.texttext-generationn<1K0 likes39 downloads5mo agoHugging Face22SAP /diaforge-utc-r-0725gated DiaFORGE UTC: Unified Tool-Calling Conversations Dataset Dataset for our paper Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky which includes 5000 enterprise tools and the corresponding dialogues generated using DiaFORGE UTC data engine. The dataset is generated with the data generation engine described in Figure 1. The engine simulates a user agent and an assistant agent in a dialogue, where the user agent has a persona and the… See the full description on the dataset page: https://huggingface.co/datasets/SAP/diaforge-utc-r-0725.texttext-generation1K<n<10K10 likes33 downloads1y agoHugging Face23sapbot /grok-4.1-fast-instruct-308xTrace of Grok 4.1 Fast LLM. WARNING: This trace was made WITHOUT reasoning. Use it to finetune only instruct models. Data count (Total: 308): English - 198 Russian - 110 Data is presented in {"messages":[{"role":"user", "content":"Prompt"}, {"role":"assistant", "content": "Response"}]} format and each conversation split by newline. This model was NOT free, and I had to use OpenRouter for it. Crypto donations for future projects like this are available on my personal page texttext-generationn<1K1 likes25 downloads5mo agoHugging Face24sapbot /gemma-4-31b-it-304xTrace of Gemma 4 31B LLM. Data count (Total: 304): English - 194 Russian - 110 Data is presented in {"messages":[{"role":"user", "content":"Prompt"}, {"role":"assistant", "content": "Response"}]} format and each conversation split by newline. texttext-generationn<1K0 likes24 downloads5mo agoHugging Face25Saptak123 /Bhagavad-Gita_Dataset Srimad Bhagavad Gita Dataset A parallel corpus of the Srimad Bhagavad Gita containing original verses in Sanskrit (sa), with Hindi (hi) and English (en) translations.This dataset is suitable for translation, text generation, and feature extraction tasks. Dataset Details Based on: This dataset is based on a book called 'Srimad Bhagavad Gita' kept at Central Archaelogical Library, New Delhi Original Book Source (IGNCA): https://ignca.gov.in/Asi_data/279.pdf… See the full description on the dataset page: https://huggingface.co/datasets/Saptak123/Bhagavad-Gita_Dataset.tabulartranslationn<1K0 likes21 downloads8mo agoHugging Face26sapbot /deepseek-v4-flash-instruct-308xTrace of DeepSeek V4 Flash LLM. WARNING: This trace was made WITHOUT reasoning. Use it to finetune only instruct models. Data count (Total: 308): English - 198 Russian - 110 Data is presented in {"messages":[{"role":"user", "content":"Prompt"}, {"role":"assistant", "content": "Response"}]} format and each conversation split by newline. This model was NOT free, and I had to use OpenRouter for it. Crypto donations for future projects like this are available on my personal page texttext-generationn<1K0 likes20 downloads5mo agoHugging Face27sapbot /gemma-3n-4b-distill-smollm2-360m-instruct-425xTrace of Gemma 3n 4B Distill SmolLM2 360M Instruct LLM by sapbot (me). Data count (Total: 425): English - 209 Russian - 216 Data is presented in ShareGPT format and each conversation split by newline. Note: This was added more as a "examples" of this model's outputs. Of course you will not distill a distilled model (I hope). Brought to you by sapbot from Romarchive texttext-generationn<1K0 likes19 downloads5mo agoHugging Face28sapbot /grok-4.1-fast-instruct-308x-cot Grok 4.1 Fast Instruct 308x traces with added russian CoT Traces generated using RU-CoT-Generator and google/gemma-3-4b-it as CoT generator. texttext-generationn<1K0 likes19 downloads3mo agoHugging Face29sapbot /coderppl CoderPPL A curated code perplexity evaluation corpus — 9,324 lines of real-world, working code across 24 files and 5 programming languages. Source: github.com/sapbotgit/code-doodles This dataset is designed to measure code perplexity (PPL) — how well a language model predicts actual hand-written code across multiple languages and programming paradigms. Contents Language Files Examples Python 5 LLM trainers, proxy scanner, fine-tuning tools JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/sapbot/coderppl.texttext-generation1K<n<10K0 likes17 downloads4mo agoHugging Face30sapbot /yandexq-qa-100 YandexQ QA (100 subset) Traces generated using RU-CoT-Generator and liquid/lfm-2.5-8b-a1b as CoT generator. Dataset based on sapbot/yandexq-qa-chatml. texttext-generationn<1K0 likes17 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.