datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-supervised-datasettinyMMLU
tinyMMLU
Welcome to tinyMMLU! This dataset serves as a concise version of the MMLU dataset, offering a subset of 100 data points selected from the original compilation.
tinyMMLU is designed to enable users to efficiently estimate the performance of a large language model (LLM) with reduced dataset size, saving computational resources
while maintaining the essence of the MMLU evaluation.
Features
Compact Dataset: With only 100 data points, tinyMMLU provides a swift… See the full description on the dataset page: https://huggingface.co/datasets/tinyBenchmarks/tinyMMLU.tiny-aya-l2-thinker-multilingual-reasoning
Tiny Aya L2 Multilingual Reasoning (44 languages)
Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker.
Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English.
Data source
Prompts from AM-DeepSeek-R1-0528-Distilled
Thinking traces and outputs distilled from gpt-oss-120b
Translated with command-a-translate and DeepSeek-V3
Languages (44)
Language
Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.tiny_qa_benchmark_pp
Tiny QA Benchmark++ (TQB++)
Tiny QA Benchmark++ (TQB++) is an ultra-lightweight evaluation suite designed to expose critical failures in Large Language Model (LLM) systems within seconds. It serves as the LLM analogue of software unit tests, ideal for rapid CI/CD checks, prompt engineering, and continuous quality assurance in modern LLMOps.
This Hugging Face dataset repository hosts the core English dataset and various synthetically generated multilingual and topical dataset packs… See the full description on the dataset page: https://huggingface.co/datasets/vincentkoc/tiny_qa_benchmark_pp.TinyTextThe entire NanoPhi Dataset is at train.jsonl
Separate Tasks Include
Math (Metamath, mammoth)
Code (Code Search Net)
Logic (Open-platypus)
Roleplay (PIPPA, RoleplayIO)
Textbooks (Tiny-text, Sciphi)
Textbook QA (Orca-text, Tiny-webtext)
tiny-aya-l2-thinker-multilingual-reasoning
Tiny Aya L2 Multilingual Reasoning (44 languages)
Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker.
Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English.
Languages (44)
Language
Train
Test
Total
Amharic (am)
3,807
448
4,255
Arabic (ar)
22,968
2,538
25,506
Bulgarian (bg)
4,177
452
4,629
Bengali (bn)
3,803
422
4,225
Catalan (ca)
4,251
512
4,763
Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.tiny-singleturn-chat-kotiny-slm-pretraining-corpus
🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB)
A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures.
100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders.
📊 Dataset Statistics
Total Documents: 20,066,075
Train: 19,663,898
Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.tinybrain-instruct-sft-200k
TinyBrain Instruct 200K
A 196k+ row English SFT dataset for training tiny instruction-following language models.
TinyBrain Instruct 200K is a synthetic supervised fine-tuning dataset made for small language models, especially models around 100M–500M parameters.
The dataset focuses on short, clear, learnable assistant responses across education, basic math reasoning, clean conversation, planning, simplification, simple coding, and honesty/uncertainty behavior.
Most… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/tinybrain-instruct-sft-200k.tiny-instruct-koVietnamese-nampdn-ai-tiny-webtext-gg-translatedtiny-llm-synthetic-qa
Tiny-LLM: Synthetic Question-Answering Dataset
Dataset Description
This dataset was created for the fine-tuning stage of the Tiny-LLM Project, a project focused on training and evaluating compact language models from scratch.
It contains 706,727 high-quality, synthetic multi-turn Question-Answering (Q&A) conversations in English, generated using the Gemini API. The dataset was designed to teach small models instruction-following capabilities across a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel8/tiny-llm-synthetic-qa.Pegasus-Tiny-250K
Pegasus-Tiny-250K
Pegasus-Tiny-250K is a compact, high-quality mathematical reasoning dataset curated by prithivMLmods and hosted on Hugging Face. It contains approximately ~291K structured reasoning traces in Parquet format, optimized for efficient training, evaluation, and reasoning-aligned fine-tuning of AI models. This dataset provides diverse mathematics-focused problem statements paired with detailed step-by-step reasoning solutions. Pegasus-Tiny-250K emphasizes clear… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Pegasus-Tiny-250K.tiny_qa_benchmark
Tiny QA Benchmark (Original English Core for TQB++)
This dataset (vincentkoc/tiny_qa_benchmark) is the original 52-item English Question-Answering set. It now serves as the immutable "gold standard" core for the expanded Tiny QA Benchmark++ (TQB++) project.
The TQB++ project builds upon this core dataset by introducing a powerful synthetic generation toolkit, pre-built multilingual datasets, and a comprehensive framework for rapid LLM smoke testing.
For the full TQB++ toolkit, the… See the full description on the dataset page: https://huggingface.co/datasets/vincentkoc/tiny_qa_benchmark.TinyGuanaco_DE
Dataset Card for TinyGuanaco_DE
TinyGuanaco_DE
is intended for development purposes: use TinyGuanaco_DE for prototyping your code
is comprised of German texts only (hence DE)
is really small: the train split has 4 instances and the test split has 2 instances
has 3 columns: index, query, and reply
the query column contains concatenations of a context ("Kontext:\n...") and a question ("Frage:\n...") that can be answered by knowing the context
the reply column contains the according… See the full description on the dataset page: https://huggingface.co/datasets/mdroth/TinyGuanaco_DE.tinysynth-reasoning
TinySynth Reasoning Primitives
Synthetic training data for teaching small language models stable state
representation and controlled reasoning operations — entity/attribute
binding, state persistence, mutation, transfer, reference resolution,
current-vs-cumulative distinctions, and claim validation — in a
systems/computing vocabulary.
Every example is generated from a hidden symbolic world and verified by a
symbolic solver before any natural language is produced:
semantic state… See the full description on the dataset page: https://huggingface.co/datasets/projenix/tinysynth-reasoning.TinySYNCR
TinySYNCR
Tiny preview subset of SYNCR with 25 samples per task.
from datasets import load_dataset
ds = load_dataset(
"CrossVideoReasoning/TinySYNCR",
split="test"
)
Each row contains:
category
task
question
options
answer
two or more video columns
License
The SYNCR dataset (including video assets and annotations) is provided under the Creative Commons Attribution 4.0 International (CC-BY 4.0) License.
The code for benchmark generation and evaluation is… See the full description on the dataset page: https://huggingface.co/datasets/CrossVideoReasoning/TinySYNCR.vietllama-tiny-envi
Instruction dataset for fine-tuning
Dataset contains original dataset [lima, orca-mini, alpaca data, alpaca finance, GPTeacher] and their Vietnamese translations
Suggested use cases: Fine-tuning Vietnamese LLM
tinymistral-hypnosis-instruct-preprocessedDataset created for accelerated processing. Embeddings from this fine model:
Locutusque/TinyMistral-248M-Instruct
TinyStories-mini-Interaction-SFT
Dataset Card for ReactiveAI/TinyStories-mini-Interaction-SFT
Interaction Supervised Fine-Tuning Dataset based on roneneldan/TinyStories dataset, made for Reactive
Transformer second training stage Proof-of-Concept.
Full version available in ReactiveAI/TinyStories-Interaction-SFT
Dataset Details
Dataset Description
Curated by: Reactive AI
Language(s) (NLP): English
License: apache-2.0
Uses
This dataset is made for Supervised… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/TinyStories-mini-Interaction-SFT.TinyStories-MRL
Dataset Card for ReactiveAI/TinyStories-MRL
Synthetic Memory Reinforcement Learning dataset for Proof-of-Concept Reactive Transformer models.
Dataset is divided into subsets, used in different Curriculum Stage of MRL training - each subset have
different number of follow-up interactions, could use different strategy, and have train and validation
splits.
After first experiments with MRL, we decided to abandon single step and two steps stages. That's because with single
step… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/TinyStories-MRL.Bilge-Tiny-MathQs
Bilge-Tiny-MathQs
This dataset contains 310 Turkish mathematics questions of varying difficulty, synthetically generated using google/gemma-4-31B-it. It primarily targets upper-primary and lower-secondary mathematics.
Each JSONL record contains the following fields:
problem: The question text
solution: A step-by-step solution
answer: The numerical answer
Gacrux-Tiny-1M
Gacrux-Tiny-1M
Gacrux-Tiny-1M is a compact, high-quality reasoning dataset curated by prithivMLmods, containing ~1.06M chain-of-thought reasoning traces optimized for mathematical problem solving, algorithmic coding challenges, and structured reasoning across competitive programming tasks. This dataset is ideal for lightweight reasoning model training and benchmarking. The dataset provides real structured problem statements with detailed reasoning step-by-step solutions that… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Gacrux-Tiny-1M.tiny-multiturn-chat-koTinyLM
TinyLM Data
This dataset in data.txt is a collection of user/ai conversations across domains such as science, math, programming and writing. It also contains general conversation data.
The dataset is designed to be used for the training of SLMs (Small Language Models).
Format
This is an example of a conversation in the dataset:
<|data|>
<|user|> What is a comet?
<|assistant|> A comet is a big ball of ice and rock. <|endoftext|>
<|user|> Does it look cool?
<|assistant|>… See the full description on the dataset page: https://huggingface.co/datasets/AGofficial/TinyLM.TinyStories-Interaction-SFT
Dataset Card for ReactiveAI/TinyStories-Interaction-SFT
Improved version of ReactiveAI/TinyStories-mini-Interaction-SFT - about 4x more rows, improved generation prompt
and additional post-processing for more diverse dataset. Includes all examples from v1 Dataset and over 75k new ones, post-processed to include more random naming.
Dataset Details
Dataset Description
Interaction Supervised Fine-Tuning Dataset based on roneneldan/TinyStories dataset, made… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/TinyStories-Interaction-SFT.Vietnamese-mabryCodes-tiny-cot-alpaca-gg-translatedtiny-think-sft-math-n-stem
Shekswess/tiny-think-sft-math-n-stem
Overview
Supervised fine-tuning (SFT) dataset built from allenai/Dolci-Think-SFT-7B plus GSM8K like think-style SFT from openai/gsm8k, using the facebook/MobileLLM-R1-140M-base tokenizer and chat template. This dataset targets math and STEM reasoning.
Dataset Details
Build date: 2026-01-10
Sources: 4
Rows: 29,149
Tokens: 59,999,048 (below budget; used all available tokens)
Max sequence length: 4096 tokens per example (chat… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/tiny-think-sft-math-n-stem.fusion-aya-math-bench
Dataset Card for Fusion Aya Math Bench
Summary
Fusion Aya Math Bench is a multilingual, olympiad-level mathematical reasoning dataset. Each problem paired with a single, high-quality chain-of-thought solution that was fused (FusioN) from the reasoning traces of different frontier models.
Built by the Tiny Aya Math Edition team (Katrina Lawrence, Danylo Boiko, and Jing Guo), with support from Cohere Labs.
Pipeline
Derived from the open-ended… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-math-edition/fusion-aya-math-bench.tiny-think-dpo-math-n-stem
Shekswess/tiny-think-dpo-math-n-stem
Overview
Direct Preference Optimization (DPO) dataset built from allenai/Dolci-Think-DPO-7B using the facebook/MobileLLM-R1-140M-base tokenizer and chat template. This dataset targets math and STEM reasoning preferences.
Dataset Details
Build date: 2026-01-10
Sources: 5
Rows: 2,861
Tokens: 9,999,795
Max sequence length: 4096 tokens per example (both chosen and rejected)
Token budget: 10,000,000 tokens (equal strategy)… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/tiny-think-dpo-math-n-stem.
