datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-supervised-datasetTinyTextThe entire NanoPhi Dataset is at train.jsonl
Separate Tasks Include
Math (Metamath, mammoth)
Code (Code Search Net)
Logic (Open-platypus)
Roleplay (PIPPA, RoleplayIO)
Textbooks (Tiny-text, Sciphi)
Textbook QA (Orca-text, Tiny-webtext)
TinyStories-Multilingual
Novelist: TinyStories Multilingual Edition
Dataset Summary
The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short, child-safe fiction designed to stress-test literary consistency, emotional warmth, and multilingual fluency in small models. Derived from the broader Novelist ecosystem, this subset focuses on narrative simplicity paired with complex moral and social themes.
The dataset contains 15,688 high-quality stories across 28 languages. Each… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/TinyStories-Multilingual.TinyBooks-QA-Chinese本数据集已停止更新,请移步https://huggingface.co/datasets/fzmnm/TinyStoriesAdv-zh
TinyBooks-QA-Chinese
Inspired by the (TinyStories)[https://arxiv.org/abs/2305.07759] paper, where a small language model exhibits strong capabilities when trained on high-quality, 🍼baby-friendly stories synthesized by AI, I present an AI-generated Encyclopedia suitable for kindergarten and grade school levels.
This AI-synthesized dataset converts classical literature into a question-answer style curriculum with… See the full description on the dataset page: https://huggingface.co/datasets/fzmnm/TinyBooks-QA-Chinese.tinybrain-instruct-sft-200k
TinyBrain Instruct 200K
A 196k+ row English SFT dataset for training tiny instruction-following language models.
TinyBrain Instruct 200K is a synthetic supervised fine-tuning dataset made for small language models, especially models around 100M–500M parameters.
The dataset focuses on short, clear, learnable assistant responses across education, basic math reasoning, clean conversation, planning, simplification, simple coding, and honesty/uncertainty behavior.
Most… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/tinybrain-instruct-sft-200k.tinybrain-pretrain-corpus-2b
TinyBrain Pretrain Corpus 2B
A mixed-source English pretraining corpus for training small language models.
TinyBrain Pretrain Corpus 2B is a mixed-source dataset built for pretraining small causal language models, especially the TinyBrain-100M Base model.
The dataset combines educational text, factual/wiki-style text, math reasoning data, Python code-summary data, clean web text, and conversation-style data. It is designed to give small models a useful general foundation… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/tinybrain-pretrain-corpus-2b.TinyEncyclopedias-Chinese本数据集已停止更新,请移步https://huggingface.co/datasets/fzmnm/TinyStoriesAdv-zh
TinyEncyclopediasChinese
Inspired by the papers (TinyStories)[https://arxiv.org/abs/2305.07759] and (Textbooks Are All You Need)[https://arxiv.org/abs/2306.11644], where a small language model exhibits strong capabilities when trained on high-quality, kid-friendly stories synthesized by AI, I present an AI-generated Encyclopedia suitable for kindergarten and grade school levels.
This dataset follows my previous… See the full description on the dataset page: https://huggingface.co/datasets/fzmnm/TinyEncyclopedias-Chinese.TinyPython
TinyPython Tasks
TinyPython is a synthetic Python dataset inspired by the idea behind TinyStories: if the data distribution is narrow, clean, and high quality, even very small language models can learn useful structure.
Instead of broad repository code or competitive-programming solutions, TinyPython focuses on short natural-language programming tasks paired with complete, typed, standalone Python functions.
The goal is to provide a compact instruction-to-code corpus for… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/TinyPython.TinyMathStories_gpt-oss-20b
TinyMathStories
A TinyStories-style corpus extended with math and lightweight reasoning.
This dataset keeps the child-level vocabulary and short narrative style of TinyStories (Microsoft Research, Eldan & Li, 2023) and mixes in basic numeracy (counting, addition/subtraction, simple equations, fractions, measurement) and short justifications—so tiny models can practice coherent English and early math/logic.
Research, generation, and curation by AlgoDriveAI.Inspired by and… See the full description on the dataset page: https://huggingface.co/datasets/AlgoDriveAI/TinyMathStories_gpt-oss-20b.Vietnamese-nampdn-ai-tiny-webtext-gg-translatedtiny-instruct-kotinyfacts
Tinyfacts
Short explanations of things, written using only about a thousand of the most common
English words — the vocabulary Randall Munroe used for Thing Explainer, itself drawn
from the xkcd comic Up Goer Five.
Writing under that constraint forces a particular kind of prose. There is no word for
photosynthesis, or gravity, or engine, so a text has to reach the idea by other
means: green things that eat light, the way everything pulls on everything else, the
part of the car… See the full description on the dataset page: https://huggingface.co/datasets/Stur86/tinyfacts.dfm11-mathagentic-tinygsm-python
dfm11-mathagentic-tinygsm-python
English arithmetic word problems converted into native Python tool-call trajectories with precomputed tool responses and boxed final answers.
Contents
Rows: 367,749
Shards: 4
Format: deterministic gzip JSON Lines in data/train-*.jsonl.gz
Schema: tools, four-message native tool trajectory, execution metadata,
stable source ID, source revision, and admission status
Intended repository:… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm11-mathagentic-tinygsm-python.Tiny-Ko-Stories
Tiny-Ko-Stories
English version is available below.
Tiny-Ko-Stories는 TinyStories에서 영감을 받은 한국어 이야기 데이터셋입니다.
TinyStories는 제한된 고품질 데이터셋을 사용하면, 소형 모델이라도 추론 능력과 창의력을 발휘할 수 있음을 보였습니다.
우리가 확인하려는 것은 단순합니다.
이 현상이 한국어에서도 재현될까?
이를 확인하려면 번역 데이터셋만으로는 부족했습니다. 한국어다운 이름, 문장 리듬, 의성어와 의태어, 색채어, 작은 사건 구조를 포함하려면 처음부터 한국어로 만든 이야기가 필요했습니다. 그래서 Tiny-Ko-Stories는 영어 TinyStories를 번역하는 대신, 한국어로 짧은 이야기를 새로 생성하고 여러 단계의 검수를 거쳐 구성했습니다.
데이터셋 요약
항목
값
레코드 수
2,003,542
형식
JSONL
공개… See the full description on the dataset page: https://huggingface.co/datasets/psymon/Tiny-Ko-Stories.tiny-vintage-completions
Tiny vintage completions
Synthetic vintage texts, with a cutoff date for year 1900.
Based on unique 2-3 word seeds, extracted from croqaz/Vintage-v1, croqaz/Vintage-v2 and Haykgrigorian/English-historical-corpus-1800-1875.
Check the files seeds1.txt and seeds2.txt.
Generated by TypeWriter-7B-base and Talkie-13B-base completions.
Citation
If you find this dataset valuable, please consider citing:
@misc{Tiny-vintage-completions,
title = {Tiny vintage completions}… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/tiny-vintage-completions.tinychat_conversations
TinyChat Conversations
A specialized conversational dataset for midtraining language models with identity awareness and personality grounding. This dataset combines model identity conversations with question-answer pairs generated from personal blog content to infuse the model with a distinct conversational style and knowledge base.
Dataset Description
This dataset contains conversational data in JSONL format, designed for midtraining chat models alongside multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/jasonacox/tinychat_conversations.TinyShell
TinyShell Dataset
TinyShell is a JSON Lines dataset of natural-language shell instructions paired
with reference commands and semantic annotations. It covers Linux, macOS, and
Windows examples across Bash, Zsh, and PowerShell-style environments.
This release contains 65,848 records in three canonical splits:
Split
Records
File
Train
46,093
data/train.jsonl
Validation
9,877
data/validation.jsonl
Test
9,878
data/test.jsonl
Release contents
The… See the full description on the dataset page: https://huggingface.co/datasets/tharunpranavsakthivel/TinyShell.SuperWiki-Tiny
Dataset Card for SuperWiki-Tiny
Waifu to catch your attention.
Dataset Details
Dataset Description
SuperWiki-Tiny is a english only subset of the SuperWikipedia-NEXT (SuperWikiNEXT-32B)
Curated by: KaraKaraWitch
Funded by: Recursal.ai (I work there lol)
Shared by: KaraKaraWitch
Language(s) (NLP): English Only.
License: cc-by-sa-4.0,
Dataset Sources
Source Data: https://dumps.wikimedia.org/other/enterprise_html/
Dataset Summary
Refer… See the full description on the dataset page: https://huggingface.co/datasets/recursal/SuperWiki-Tiny.bun-server-bench-trajectories
bun-server-bench trajectories
Supervised fine-tuning and patch trajectories exported from bun-server-bench,
a benchmark for evaluating coding agents on real-world Bun server engineering tasks.
Every record comes from an agent run that passed both the public and hidden tests
for its task — these are verified solutions, not raw attempts. The benchmark engineers
each task so that a plausible-but-wrong implementation passes the visible tests and
fails the hidden ones, so a passing… See the full description on the dataset page: https://huggingface.co/datasets/tinycomputerai/bun-server-bench-trajectories.xsum_tinyThis dataset is a subset of https://huggingface.co/datasets/EdinburghNLP/xsum.
The training set is composed of 2,000 examples of the original training set and the test set is composed of 1,000 examples of the original validation set.
starhopp3r-TinyChat
starhopp3r/TinyChat
This is an unofficial reformatted version of starhopp3r/TinyChat.
It contains about 1 million conversations generated by GPT-4o mini using basic English.
The conversations have been cleaned and put into ShareGPT-like format.
Duplicates have been removed.
All credit belongs to the original author.
Footnote: These conversations have a strong mono no aware feeling in my opinion.
tiny-multiturn-chat-koTinyFinewebEdu-kofineweb-edu에서 int_score가 4 이상인 데이터만 필터링한 후 DeepSeek-v3을 이용해 데이터를 단순한 형태의 문장과 간단한 어휘로 구성되도록 변환한 데이터입니다.
비용 문제로 데이터는 67k 개만 있습니다.
TinyLLMPretrainingCore
Synthetic Simple-English Subject Explanations Dataset
Dataset Summary
This dataset contains synthetic, GPT-generated texts that explain a wide range of subjects using simple English.Each subject is expanded into multiple long-form explanations that repeat key ideas across different styles, perspectives, and framing strategies.
The dataset is designed to emphasize clarity, redundancy, and consistency, making it useful for educational NLP, simplification tasks, and… See the full description on the dataset page: https://huggingface.co/datasets/MaxHastings/TinyLLMPretrainingCore.TinyChat-ITA
TinyChat-ITA
Questo dataset fornisce coppie domanda-risposta in lingua italiana, pensate per applicazioni di chatbot e modelli conversazionali. Ogni entry contiene:
Una domanda breve e naturale (input).
Una risposta chiara e coerente (response).
Tutti i dati sono memorizzati in formato JSONL, dove ogni riga rappresenta un oggetto JSON valido. Le risposte non parsabili sono state salvate separatamente per la revisione.
Curato da: Mattimax per M.INC. (M.INC. profile)
Condiviso… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/TinyChat-ITA.cnn_dailymail_tinyThis dataset is a subset of https://huggingface.co/datasets/cnn_dailymail.
The training set is composed of 2,000 examples of the original training set and the test set is composed of 1,000 examples of the original validation set.
We use the version 1.0.0 of the CNN/DailyMail dataset.
tinypairs
TinyPairs
TinyPairs is a dataset of 1000 preprocessed input-target pairs derived from roneneldan/TinyStories.This dataset is formatted as a simple JSON file for easy use in training small-scale language models.
📜 Format: Each entry consists of:
{
"input": "Sue liked to study. She would sit upstairs in her room and look at her books.",
"target": "Sue had a dog named Max. Max was deaf, but he was a good dog."
}
🔧 How It Was Generated
The dataset was extracted and… See the full description on the dataset page: https://huggingface.co/datasets/teleprint-me/tinypairs.The-Tiny-Purr-2
purrgpt-community/The-Tiny-Purr-2
The Tiny Purr is a turn base dataset that contains different lengths of conversation!
What is cant do:
Use tools
Do web search (it can, hopefully)
For what is:
For a very frendly cat AI chatbot
WARNING: There might be some mistakes
adgen_tinyThis dataset is a subset of the advertising generation dataset proposed by https://aclanthology.org/D19-1321/.
The training set is composed of 2,000 examples of the original training set and the test set is composed of 1,000 examples of the original validation set.
TinyJenna-Uncensored-v01
Uncensored Alpaca Dataset: A New Frontier in Language Models
This dataset is a collection of uncensored prompts and responses in the Alpaca format. It aims to provide a diverse and unfiltered source of data for training language models, pushing the boundaries of what these models can understand and generate.
What Makes This Dataset Different?
Uncensored: This dataset includes prompts and responses that touch upon topics that are often censored or avoided in traditional datasets.… See the full description on the dataset page: https://huggingface.co/datasets/V3N0M/TinyJenna-Uncensored-v01.
