CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes45k downloads2y agoHugging Face02tinyBenchmarks /tinyMMLU tinyMMLU Welcome to tinyMMLU! This dataset serves as a concise version of the MMLU dataset, offering a subset of 100 data points selected from the original compilation. tinyMMLU is designed to enable users to efficiently estimate the performance of a large language model (LLM) with reduced dataset size, saving computational resources while maintaining the essence of the MMLU evaluation. Features Compact Dataset: With only 100 data points, tinyMMLU provides a swift… See the full description on the dataset page: https://huggingface.co/datasets/tinyBenchmarks/tinyMMLU.textquestion-answeringn<1K24 likes5.1k downloads2y agoHugging Face03CohereLabs /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Data source Prompts from AM-DeepSeek-R1-0528-Distilled Thinking traces and outputs distilled from gpt-oss-120b Translated with command-a-translate and DeepSeek-V3 Languages (44) Language Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M6 likes1.2k downloads13d agoHugging Face04vincentkoc /tiny_qa_benchmark_pp Tiny QA Benchmark++ (TQB++) Tiny QA Benchmark++ (TQB++) is an ultra-lightweight evaluation suite designed to expose critical failures in Large Language Model (LLM) systems within seconds. It serves as the LLM analogue of software unit tests, ideal for rapid CI/CD checks, prompt engineering, and continuous quality assurance in modern LLMOps. This Hugging Face dataset repository hosts the core English dataset and various synthetically generated multilingual and topical dataset packs… See the full description on the dataset page: https://huggingface.co/datasets/vincentkoc/tiny_qa_benchmark_pp.textquestion-answeringn<1K3 likes490 downloads1y agoHugging Face05VatsaDev /TinyTextThe entire NanoPhi Dataset is at train.jsonl Separate Tasks Include Math (Metamath, mammoth) Code (Code Search Net) Logic (Open-platypus) Roleplay (PIPPA, RoleplayIO) Textbooks (Tiny-text, Sciphi) Textbook QA (Orca-text, Tiny-webtext) textquestion-answering1M<n<10M34 likes407 downloads2y agoHugging Face06erenyeager-1 /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Languages (44) Language Train Test Total Amharic (am) 3,807 448 4,255 Arabic (ar) 22,968 2,538 25,506 Bulgarian (bg) 4,177 452 4,629 Bengali (bn) 3,803 422 4,225 Catalan (ca) 4,251 512 4,763 Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M0 likes220 downloads14d agoHugging Face07OpenLab-NLP /tiny-singleturn-chat-kotextquestion-answering10K<n<100K0 likes201 downloads10mo agoHugging Face08JustACluelessKidAtSchool /tiny-slm-pretraining-corpus 🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB) A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures. 100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders. 📊 Dataset Statistics Total Documents: 20,066,075 Train: 19,663,898 Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.tabulartext-generation10M<n<100M0 likes198 downloads1mo agoHugging Face09exnivo /tinybrain-instruct-sft-200k TinyBrain Instruct 200K A 196k+ row English SFT dataset for training tiny instruction-following language models. TinyBrain Instruct 200K is a synthetic supervised fine-tuning dataset made for small language models, especially models around 100M–500M parameters. The dataset focuses on short, clear, learnable assistant responses across education, basic math reasoning, clean conversation, planning, simplification, simple coding, and honesty/uncertainty behavior. Most… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/tinybrain-instruct-sft-200k.texttext-generation100K<n<1M3 likes190 downloads3mo agoHugging Face10OpenLab-NLP /tiny-instruct-kotextquestion-answering10K<n<100K1 likes126 downloads9mo agoHugging Face115CD-AI /Vietnamese-nampdn-ai-tiny-webtext-gg-translatedtextquestion-answering1M<n<10M10 likes98 downloads3y agoHugging Face12Gabriel8 /tiny-llm-synthetic-qa Tiny-LLM: Synthetic Question-Answering Dataset Dataset Description This dataset was created for the fine-tuning stage of the Tiny-LLM Project, a project focused on training and evaluating compact language models from scratch. It contains 706,727 high-quality, synthetic multi-turn Question-Answering (Q&A) conversations in English, generated using the Gemini API. The dataset was designed to teach small models instruction-following capabilities across a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel8/tiny-llm-synthetic-qa.textquestion-answering100K<n<1M2 likes98 downloads11mo agoHugging Face13prithivMLmods /Pegasus-Tiny-250K Pegasus-Tiny-250K Pegasus-Tiny-250K is a compact, high-quality mathematical reasoning dataset curated by prithivMLmods and hosted on Hugging Face. It contains approximately ~291K structured reasoning traces in Parquet format, optimized for efficient training, evaluation, and reasoning-aligned fine-tuning of AI models. This dataset provides diverse mathematics-focused problem statements paired with detailed step-by-step reasoning solutions. Pegasus-Tiny-250K emphasizes clear… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Pegasus-Tiny-250K.texttext-generation100K<n<1M2 likes92 downloads10mo agoHugging Face14vincentkoc /tiny_qa_benchmark Tiny QA Benchmark (Original English Core for TQB++) This dataset (vincentkoc/tiny_qa_benchmark) is the original 52-item English Question-Answering set. It now serves as the immutable "gold standard" core for the expanded Tiny QA Benchmark++ (TQB++) project. The TQB++ project builds upon this core dataset by introducing a powerful synthetic generation toolkit, pre-built multilingual datasets, and a comprehensive framework for rapid LLM smoke testing. For the full TQB++ toolkit, the… See the full description on the dataset page: https://huggingface.co/datasets/vincentkoc/tiny_qa_benchmark.textquestion-answeringn<1K1 likes75 downloads1y agoHugging Face15mdroth /TinyGuanaco_DE Dataset Card for TinyGuanaco_DE TinyGuanaco_DE is intended for development purposes: use TinyGuanaco_DE for prototyping your code is comprised of German texts only (hence DE) is really small: the train split has 4 instances and the test split has 2 instances has 3 columns: index, query, and reply the query column contains concatenations of a context ("Kontext:\n...") and a question ("Frage:\n...") that can be answered by knowing the context the reply column contains the according… See the full description on the dataset page: https://huggingface.co/datasets/mdroth/TinyGuanaco_DE.textquestion-answeringn<1K0 likes71 downloads3y agoHugging Face16projenix /tinysynth-reasoning TinySynth Reasoning Primitives Synthetic training data for teaching small language models stable state representation and controlled reasoning operations — entity/attribute binding, state persistence, mutation, transfer, reference resolution, current-vs-cumulative distinctions, and claim validation — in a systems/computing vocabulary. Every example is generated from a hidden symbolic world and verified by a symbolic solver before any natural language is produced: semantic state… See the full description on the dataset page: https://huggingface.co/datasets/projenix/tinysynth-reasoning.tabulartext-generation1M<n<10M0 likes63 downloads9d agoHugging Face17CrossVideoReasoning /TinySYNCR TinySYNCR Tiny preview subset of SYNCR with 25 samples per task. from datasets import load_dataset ds = load_dataset( "CrossVideoReasoning/TinySYNCR", split="test" ) Each row contains: category task question options answer two or more video columns License The SYNCR dataset (including video assets and annotations) is provided under the Creative Commons Attribution 4.0 International (CC-BY 4.0) License. The code for benchmark generation and evaluation is… See the full description on the dataset page: https://huggingface.co/datasets/CrossVideoReasoning/TinySYNCR.textvideo-classificationn<1K0 likes60 downloads5mo agoHugging Face18infCapital /vietllama-tiny-envi Instruction dataset for fine-tuning Dataset contains original dataset [lima, orca-mini, alpaca data, alpaca finance, GPTeacher] and their Vietnamese translations Suggested use cases: Fine-tuning Vietnamese LLM textquestion-answering100K<n<1M0 likes58 downloads3y agoHugging Face19jtatman /tinymistral-hypnosis-instruct-preprocessedDataset created for accelerated processing. Embeddings from this fine model: Locutusque/TinyMistral-248M-Instruct textquestion-answering1M<n<10M3 likes58 downloads3y agoHugging Face20ReactiveAI /TinyStories-mini-Interaction-SFT Dataset Card for ReactiveAI/TinyStories-mini-Interaction-SFT Interaction Supervised Fine-Tuning Dataset based on roneneldan/TinyStories dataset, made for Reactive Transformer second training stage Proof-of-Concept. Full version available in ReactiveAI/TinyStories-Interaction-SFT Dataset Details Dataset Description Curated by: Reactive AI Language(s) (NLP): English License: apache-2.0 Uses This dataset is made for Supervised… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/TinyStories-mini-Interaction-SFT.textquestion-answering10K<n<100K0 likes58 downloads1y agoHugging Face21ReactiveAI /TinyStories-MRL Dataset Card for ReactiveAI/TinyStories-MRL Synthetic Memory Reinforcement Learning dataset for Proof-of-Concept Reactive Transformer models. Dataset is divided into subsets, used in different Curriculum Stage of MRL training - each subset have different number of follow-up interactions, could use different strategy, and have train and validation splits. After first experiments with MRL, we decided to abandon single step and two steps stages. That's because with single step… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/TinyStories-MRL.textreinforcement-learning10K<n<100K0 likes56 downloads1y agoHugging Face22BILGEM-AI /Bilge-Tiny-MathQs Bilge-Tiny-MathQs This dataset contains 310 Turkish mathematics questions of varying difficulty, synthetically generated using google/gemma-4-31B-it. It primarily targets upper-primary and lower-secondary mathematics. Each JSONL record contains the following fields: problem: The question text solution: A step-by-step solution answer: The numerical answer textquestion-answeringn<1K2 likes54 downloads1mo agoHugging Face23prithivMLmods /Gacrux-Tiny-1M Gacrux-Tiny-1M Gacrux-Tiny-1M is a compact, high-quality reasoning dataset curated by prithivMLmods, containing ~1.06M chain-of-thought reasoning traces optimized for mathematical problem solving, algorithmic coding challenges, and structured reasoning across competitive programming tasks. This dataset is ideal for lightweight reasoning model training and benchmarking. The dataset provides real structured problem statements with detailed reasoning step-by-step solutions that… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Gacrux-Tiny-1M.texttext-generation1M<n<10M4 likes50 downloads7mo agoHugging Face24OpenLab-NLP /tiny-multiturn-chat-kotextquestion-answering1M<n<10M0 likes39 downloads10mo agoHugging Face25AGofficial /TinyLM TinyLM Data This dataset in data.txt is a collection of user/ai conversations across domains such as science, math, programming and writing. It also contains general conversation data. The dataset is designed to be used for the training of SLMs (Small Language Models). Format This is an example of a conversation in the dataset: <|data|> <|user|> What is a comet? <|assistant|> A comet is a big ball of ice and rock. <|endoftext|> <|user|> Does it look cool? <|assistant|>… See the full description on the dataset page: https://huggingface.co/datasets/AGofficial/TinyLM.textquestion-answering10K<n<100K1 likes38 downloads8mo agoHugging Face26ReactiveAI /TinyStories-Interaction-SFT Dataset Card for ReactiveAI/TinyStories-Interaction-SFT Improved version of ReactiveAI/TinyStories-mini-Interaction-SFT - about 4x more rows, improved generation prompt and additional post-processing for more diverse dataset. Includes all examples from v1 Dataset and over 75k new ones, post-processed to include more random naming. Dataset Details Dataset Description Interaction Supervised Fine-Tuning Dataset based on roneneldan/TinyStories dataset, made… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/TinyStories-Interaction-SFT.texttext-generation100K<n<1M1 likes35 downloads1y agoHugging Face275CD-AI /Vietnamese-mabryCodes-tiny-cot-alpaca-gg-translatedtextquestion-answering100K<n<1M23 likes32 downloads3y agoHugging Face28Shekswess /tiny-think-sft-math-n-stem Shekswess/tiny-think-sft-math-n-stem Overview Supervised fine-tuning (SFT) dataset built from allenai/Dolci-Think-SFT-7B plus GSM8K like think-style SFT from openai/gsm8k, using the facebook/MobileLLM-R1-140M-base tokenizer and chat template. This dataset targets math and STEM reasoning. Dataset Details Build date: 2026-01-10 Sources: 4 Rows: 29,149 Tokens: 59,999,048 (below budget; used all available tokens) Max sequence length: 4096 tokens per example (chat… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/tiny-think-sft-math-n-stem.textquestion-answering10K<n<100K0 likes32 downloads8mo agoHugging Face29tiny-aya-math-edition /fusion-aya-math-bench Dataset Card for Fusion Aya Math Bench Summary Fusion Aya Math Bench is a multilingual, olympiad-level mathematical reasoning dataset. Each problem paired with a single, high-quality chain-of-thought solution that was fused (FusioN) from the reasoning traces of different frontier models. Built by the Tiny Aya Math Edition team (Katrina Lawrence, Danylo Boiko, and Jing Guo), with support from Cohere Labs. Pipeline Derived from the open-ended… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-math-edition/fusion-aya-math-bench.texttext-generation1K<n<10K2 likes32 downloads3mo agoHugging Face30Shekswess /tiny-think-dpo-math-n-stem Shekswess/tiny-think-dpo-math-n-stem Overview Direct Preference Optimization (DPO) dataset built from allenai/Dolci-Think-DPO-7B using the facebook/MobileLLM-R1-140M-base tokenizer and chat template. This dataset targets math and STEM reasoning preferences. Dataset Details Build date: 2026-01-10 Sources: 5 Rows: 2,861 Tokens: 9,999,795 Max sequence length: 4096 tokens per example (both chosen and rejected) Token budget: 10,000,000 tokens (equal strategy)… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/tiny-think-dpo-math-n-stem.textquestion-answering1K<n<10K0 likes30 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.