CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes45k downloads2y agoHugging Face02tinyBenchmarks /tinyMMLU tinyMMLU Welcome to tinyMMLU! This dataset serves as a concise version of the MMLU dataset, offering a subset of 100 data points selected from the original compilation. tinyMMLU is designed to enable users to efficiently estimate the performance of a large language model (LLM) with reduced dataset size, saving computational resources while maintaining the essence of the MMLU evaluation. Features Compact Dataset: With only 100 data points, tinyMMLU provides a swift… See the full description on the dataset page: https://huggingface.co/datasets/tinyBenchmarks/tinyMMLU.textquestion-answeringn<1K24 likes4.9k downloads2y agoHugging Face03tinyBenchmarks /tinyAI2_arc tinyAI2_arc Welcome to tinyAI2_arc! This dataset serves as a concise version of the AI2_arc challenge dataset, offering a subset of 100 data points selected from the original compilation. tinyAI2_arc is designed to enable users to efficiently estimate the performance of a large language model (LLM) with reduced dataset size, saving computational resources while maintaining the essence of the ARC challenge evaluation. Features Compact Dataset: With only 100 data… See the full description on the dataset page: https://huggingface.co/datasets/tinyBenchmarks/tinyAI2_arc.question-answeringn<1K4 likes2.6k downloads2y agoHugging Face04tinyBenchmarks /tinyTruthfulQA tinyTruthfulQA Welcome to tinyTruthfulQA! This dataset serves as a concise version of the truthfulQA dataset, offering a subset of 100 data points selected from the original compilation. tinyTruthfulQA is designed to enable users to efficiently estimate the performance of a large language model (LLM) with reduced dataset size, saving computational resources while maintaining the essence of the truthfulQA evaluation. Features Compact Dataset: With only 100 data… See the full description on the dataset page: https://huggingface.co/datasets/tinyBenchmarks/tinyTruthfulQA.multiple-choicen<1K4 likes1.9k downloads2y agoHugging Face05CohereLabs /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Data source Prompts from AM-DeepSeek-R1-0528-Distilled Thinking traces and outputs distilled from gpt-oss-120b Translated with command-a-translate and DeepSeek-V3 Languages (44) Language Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M6 likes1.2k downloads15d agoHugging Face06vincentkoc /tiny_qa_benchmark_pp Tiny QA Benchmark++ (TQB++) Tiny QA Benchmark++ (TQB++) is an ultra-lightweight evaluation suite designed to expose critical failures in Large Language Model (LLM) systems within seconds. It serves as the LLM analogue of software unit tests, ideal for rapid CI/CD checks, prompt engineering, and continuous quality assurance in modern LLMOps. This Hugging Face dataset repository hosts the core English dataset and various synthetically generated multilingual and topical dataset packs… See the full description on the dataset page: https://huggingface.co/datasets/vincentkoc/tiny_qa_benchmark_pp.textquestion-answeringn<1K3 likes528 downloads1y agoHugging Face07VatsaDev /TinyTextThe entire NanoPhi Dataset is at train.jsonl Separate Tasks Include Math (Metamath, mammoth) Code (Code Search Net) Logic (Open-platypus) Roleplay (PIPPA, RoleplayIO) Textbooks (Tiny-text, Sciphi) Textbook QA (Orca-text, Tiny-webtext) textquestion-answering1M<n<10M34 likes364 downloads2y agoHugging Face08erenyeager-1 /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Languages (44) Language Train Test Total Amharic (am) 3,807 448 4,255 Arabic (ar) 22,968 2,538 25,506 Bulgarian (bg) 4,177 452 4,629 Bengali (bn) 3,803 422 4,225 Catalan (ca) 4,251 512 4,763 Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M0 likes323 downloads16d agoHugging Face09AxiomicLabs /Tiny_Theory_of_Mind Tiny Theory of Mind Tiny Theory of Mind is our first attempt at evaluating theory of mind capabilities in small language models. The benchmark covers a wide variety of ToM topics, ranging in difficulties that, for humans, would be appropriate for Pre-K through 6th grade. The benchmark is designed primarily for base-model continuation log-likelihood scoring. It does not require instruction following, chain-of-thought, or generated explanations. Random-choice accuracy is 25%.… See the full description on the dataset page: https://huggingface.co/datasets/AxiomicLabs/Tiny_Theory_of_Mind.textquestion-answering1K<n<10K24 likes241 downloads3d agoHugging Face10MedMLLM-attack /3MAD-Tiny-1Kimage-feature-extraction4 likes228 downloads2y agoHugging Face11OpenLab-NLP /tiny-singleturn-chat-kotextquestion-answering10K<n<100K0 likes202 downloads10mo agoHugging Face12exnivo /tinybrain-instruct-sft-200k TinyBrain Instruct 200K A 196k+ row English SFT dataset for training tiny instruction-following language models. TinyBrain Instruct 200K is a synthetic supervised fine-tuning dataset made for small language models, especially models around 100M–500M parameters. The dataset focuses on short, clear, learnable assistant responses across education, basic math reasoning, clean conversation, planning, simplification, simple coding, and honesty/uncertainty behavior. Most… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/tinybrain-instruct-sft-200k.texttext-generation100K<n<1M3 likes178 downloads3mo agoHugging Face135CD-AI /Vietnamese-nampdn-ai-tiny-webtext-gg-translatedtextquestion-answering1M<n<10M10 likes96 downloads3y agoHugging Face14OpenLab-NLP /tiny-instruct-kotextquestion-answering10K<n<100K1 likes93 downloads9mo agoHugging Face15prithivMLmods /Pegasus-Tiny-250K Pegasus-Tiny-250K Pegasus-Tiny-250K is a compact, high-quality mathematical reasoning dataset curated by prithivMLmods and hosted on Hugging Face. It contains approximately ~291K structured reasoning traces in Parquet format, optimized for efficient training, evaluation, and reasoning-aligned fine-tuning of AI models. This dataset provides diverse mathematics-focused problem statements paired with detailed step-by-step reasoning solutions. Pegasus-Tiny-250K emphasizes clear… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Pegasus-Tiny-250K.texttext-generation100K<n<1M2 likes92 downloads10mo agoHugging Face16Gabriel8 /tiny-llm-synthetic-qa Tiny-LLM: Synthetic Question-Answering Dataset Dataset Description This dataset was created for the fine-tuning stage of the Tiny-LLM Project, a project focused on training and evaluating compact language models from scratch. It contains 706,727 high-quality, synthetic multi-turn Question-Answering (Q&A) conversations in English, generated using the Gemini API. The dataset was designed to teach small models instruction-following capabilities across a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel8/tiny-llm-synthetic-qa.textquestion-answering100K<n<1M2 likes89 downloads1y agoHugging Face17JustACluelessKidAtSchool /tiny-slm-pretraining-corpus 🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB) A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures. 100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders. 📊 Dataset Statistics Total Documents: 20,066,075 Train: 19,663,898 Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.tabulartext-generation10M<n<100M0 likes86 downloads1mo agoHugging Face18mdroth /TinyGuanaco_DE Dataset Card for TinyGuanaco_DE TinyGuanaco_DE is intended for development purposes: use TinyGuanaco_DE for prototyping your code is comprised of German texts only (hence DE) is really small: the train split has 4 instances and the test split has 2 instances has 3 columns: index, query, and reply the query column contains concatenations of a context ("Kontext:\n...") and a question ("Frage:\n...") that can be answered by knowing the context the reply column contains the according… See the full description on the dataset page: https://huggingface.co/datasets/mdroth/TinyGuanaco_DE.textquestion-answeringn<1K0 likes74 downloads3y agoHugging Face19CrossVideoReasoning /TinySYNCR TinySYNCR Tiny preview subset of SYNCR with 25 samples per task. from datasets import load_dataset ds = load_dataset( "CrossVideoReasoning/TinySYNCR", split="test" ) Each row contains: category task question options answer two or more video columns License The SYNCR dataset (including video assets and annotations) is provided under the Creative Commons Attribution 4.0 International (CC-BY 4.0) License. The code for benchmark generation and evaluation is… See the full description on the dataset page: https://huggingface.co/datasets/CrossVideoReasoning/TinySYNCR.textvideo-classificationn<1K0 likes72 downloads5mo agoHugging Face20vincentkoc /tiny_qa_benchmark Tiny QA Benchmark (Original English Core for TQB++) This dataset (vincentkoc/tiny_qa_benchmark) is the original 52-item English Question-Answering set. It now serves as the immutable "gold standard" core for the expanded Tiny QA Benchmark++ (TQB++) project. The TQB++ project builds upon this core dataset by introducing a powerful synthetic generation toolkit, pre-built multilingual datasets, and a comprehensive framework for rapid LLM smoke testing. For the full TQB++ toolkit, the… See the full description on the dataset page: https://huggingface.co/datasets/vincentkoc/tiny_qa_benchmark.textquestion-answeringn<1K1 likes69 downloads1y agoHugging Face21ReactiveAI /TinyStories-mini-Interaction-SFT Dataset Card for ReactiveAI/TinyStories-mini-Interaction-SFT Interaction Supervised Fine-Tuning Dataset based on roneneldan/TinyStories dataset, made for Reactive Transformer second training stage Proof-of-Concept. Full version available in ReactiveAI/TinyStories-Interaction-SFT Dataset Details Dataset Description Curated by: Reactive AI Language(s) (NLP): English License: apache-2.0 Uses This dataset is made for Supervised… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/TinyStories-mini-Interaction-SFT.textquestion-answering10K<n<100K0 likes65 downloads1y agoHugging Face22projenix /tinysynth-reasoning TinySynth Reasoning Primitives Synthetic training data for teaching small language models stable state representation and controlled reasoning operations — entity/attribute binding, state persistence, mutation, transfer, reference resolution, current-vs-cumulative distinctions, and claim validation — in a systems/computing vocabulary. Every example is generated from a hidden symbolic world and verified by a symbolic solver before any natural language is produced: semantic state… See the full description on the dataset page: https://huggingface.co/datasets/projenix/tinysynth-reasoning.tabulartext-generation1M<n<10M0 likes64 downloads12d agoHugging Face23jtatman /tinymistral-hypnosis-instruct-preprocessedDataset created for accelerated processing. Embeddings from this fine model: Locutusque/TinyMistral-248M-Instruct textquestion-answering1M<n<10M3 likes62 downloads3y agoHugging Face24infCapital /vietllama-tiny-envi Instruction dataset for fine-tuning Dataset contains original dataset [lima, orca-mini, alpaca data, alpaca finance, GPTeacher] and their Vietnamese translations Suggested use cases: Fine-tuning Vietnamese LLM textquestion-answering100K<n<1M0 likes58 downloads3y agoHugging Face25ReactiveAI /TinyStories-MRL Dataset Card for ReactiveAI/TinyStories-MRL Synthetic Memory Reinforcement Learning dataset for Proof-of-Concept Reactive Transformer models. Dataset is divided into subsets, used in different Curriculum Stage of MRL training - each subset have different number of follow-up interactions, could use different strategy, and have train and validation splits. After first experiments with MRL, we decided to abandon single step and two steps stages. That's because with single step… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/TinyStories-MRL.textreinforcement-learning10K<n<100K0 likes58 downloads1y agoHugging Face26BILGEM-AI /Bilge-Tiny-MathQs Bilge-Tiny-MathQs This dataset contains 310 Turkish mathematics questions of varying difficulty, synthetically generated using google/gemma-4-31B-it. It primarily targets upper-primary and lower-secondary mathematics. Each JSONL record contains the following fields: problem: The question text solution: A step-by-step solution answer: The numerical answer textquestion-answeringn<1K2 likes49 downloads1mo agoHugging Face27prithivMLmods /Gacrux-Tiny-1M Gacrux-Tiny-1M Gacrux-Tiny-1M is a compact, high-quality reasoning dataset curated by prithivMLmods, containing ~1.06M chain-of-thought reasoning traces optimized for mathematical problem solving, algorithmic coding challenges, and structured reasoning across competitive programming tasks. This dataset is ideal for lightweight reasoning model training and benchmarking. The dataset provides real structured problem statements with detailed reasoning step-by-step solutions that… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Gacrux-Tiny-1M.texttext-generation1M<n<10M4 likes48 downloads7mo agoHugging Face28ReactiveAI /TinyStories-Interaction-SFT Dataset Card for ReactiveAI/TinyStories-Interaction-SFT Improved version of ReactiveAI/TinyStories-mini-Interaction-SFT - about 4x more rows, improved generation prompt and additional post-processing for more diverse dataset. Includes all examples from v1 Dataset and over 75k new ones, post-processed to include more random naming. Dataset Details Dataset Description Interaction Supervised Fine-Tuning Dataset based on roneneldan/TinyStories dataset, made… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/TinyStories-Interaction-SFT.texttext-generation100K<n<1M1 likes41 downloads1y agoHugging Face29Shekswess /tiny-think-sft-math-n-stem Shekswess/tiny-think-sft-math-n-stem Overview Supervised fine-tuning (SFT) dataset built from allenai/Dolci-Think-SFT-7B plus GSM8K like think-style SFT from openai/gsm8k, using the facebook/MobileLLM-R1-140M-base tokenizer and chat template. This dataset targets math and STEM reasoning. Dataset Details Build date: 2026-01-10 Sources: 4 Rows: 29,149 Tokens: 59,999,048 (below budget; used all available tokens) Max sequence length: 4096 tokens per example (chat… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/tiny-think-sft-math-n-stem.textquestion-answering10K<n<100K0 likes36 downloads8mo agoHugging Face30Shekswess /tiny-think-dpo-math-n-stem Shekswess/tiny-think-dpo-math-n-stem Overview Direct Preference Optimization (DPO) dataset built from allenai/Dolci-Think-DPO-7B using the facebook/MobileLLM-R1-140M-base tokenizer and chat template. This dataset targets math and STEM reasoning preferences. Dataset Details Build date: 2026-01-10 Sources: 5 Rows: 2,861 Tokens: 9,999,795 Max sequence length: 4096 tokens per example (both chosen and rejected) Token budget: 10,000,000 tokens (equal strategy)… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/tiny-think-dpo-math-n-stem.textquestion-answering1K<n<10K0 likes33 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.