CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes45k downloads2y agoHugging Face02VatsaDev /TinyTextThe entire NanoPhi Dataset is at train.jsonl Separate Tasks Include Math (Metamath, mammoth) Code (Code Search Net) Logic (Open-platypus) Roleplay (PIPPA, RoleplayIO) Textbooks (Tiny-text, Sciphi) Textbook QA (Orca-text, Tiny-webtext) textquestion-answering1M<n<10M34 likes364 downloads2y agoHugging Face03Dxniz /TinyStories-Multilingual Novelist: TinyStories Multilingual Edition Dataset Summary The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short, child-safe fiction designed to stress-test literary consistency, emotional warmth, and multilingual fluency in small models. Derived from the broader Novelist ecosystem, this subset focuses on narrative simplicity paired with complex moral and social themes. The dataset contains 15,688 high-quality stories across 28 languages. Each… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/TinyStories-Multilingual.texttext-generation10K<n<100K1 likes334 downloads6mo agoHugging Face04fzmnm /TinyBooks-QA-Chinese本数据集已停止更新,请移步https://huggingface.co/datasets/fzmnm/TinyStoriesAdv-zh TinyBooks-QA-Chinese Inspired by the (TinyStories)[https://arxiv.org/abs/2305.07759] paper, where a small language model exhibits strong capabilities when trained on high-quality, 🍼baby-friendly stories synthesized by AI, I present an AI-generated Encyclopedia suitable for kindergarten and grade school levels. This AI-synthesized dataset converts classical literature into a question-answer style curriculum with… See the full description on the dataset page: https://huggingface.co/datasets/fzmnm/TinyBooks-QA-Chinese.texttext-generation1K<n<10K8 likes317 downloads2y agoHugging Face05exnivo /tinybrain-instruct-sft-200k TinyBrain Instruct 200K A 196k+ row English SFT dataset for training tiny instruction-following language models. TinyBrain Instruct 200K is a synthetic supervised fine-tuning dataset made for small language models, especially models around 100M–500M parameters. The dataset focuses on short, clear, learnable assistant responses across education, basic math reasoning, clean conversation, planning, simplification, simple coding, and honesty/uncertainty behavior. Most… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/tinybrain-instruct-sft-200k.texttext-generation100K<n<1M3 likes178 downloads3mo agoHugging Face06exnivo /tinybrain-pretrain-corpus-2b TinyBrain Pretrain Corpus 2B A mixed-source English pretraining corpus for training small language models. TinyBrain Pretrain Corpus 2B is a mixed-source dataset built for pretraining small causal language models, especially the TinyBrain-100M Base model. The dataset combines educational text, factual/wiki-style text, math reasoning data, Python code-summary data, clean web text, and conversation-style data. It is designed to give small models a useful general foundation… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/tinybrain-pretrain-corpus-2b.texttext-generation1M<n<10M1 likes132 downloads3mo agoHugging Face07fzmnm /TinyEncyclopedias-Chinese本数据集已停止更新,请移步https://huggingface.co/datasets/fzmnm/TinyStoriesAdv-zh TinyEncyclopediasChinese Inspired by the papers (TinyStories)[https://arxiv.org/abs/2305.07759] and (Textbooks Are All You Need)[https://arxiv.org/abs/2306.11644], where a small language model exhibits strong capabilities when trained on high-quality, kid-friendly stories synthesized by AI, I present an AI-generated Encyclopedia suitable for kindergarten and grade school levels. This dataset follows my previous… See the full description on the dataset page: https://huggingface.co/datasets/fzmnm/TinyEncyclopedias-Chinese.texttext-generation10K<n<100K1 likes129 downloads2y agoHugging Face08BertilBraun /TinyPython TinyPython Tasks TinyPython is a synthetic Python dataset inspired by the idea behind TinyStories: if the data distribution is narrow, clean, and high quality, even very small language models can learn useful structure. Instead of broad repository code or competitive-programming solutions, TinyPython focuses on short natural-language programming tasks paired with complete, typed, standalone Python functions. The goal is to provide a compact instruction-to-code corpus for… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/TinyPython.texttext-generation1M<n<10M1 likes106 downloads3mo agoHugging Face09AlgoDriveAI /TinyMathStories_gpt-oss-20b TinyMathStories A TinyStories-style corpus extended with math and lightweight reasoning. This dataset keeps the child-level vocabulary and short narrative style of TinyStories (Microsoft Research, Eldan & Li, 2023) and mixes in basic numeracy (counting, addition/subtraction, simple equations, fractions, measurement) and short justifications—so tiny models can practice coherent English and early math/logic. Research, generation, and curation by AlgoDriveAI.Inspired by and… See the full description on the dataset page: https://huggingface.co/datasets/AlgoDriveAI/TinyMathStories_gpt-oss-20b.texttext-generation100K<n<1M0 likes99 downloads9mo agoHugging Face105CD-AI /Vietnamese-nampdn-ai-tiny-webtext-gg-translatedtextquestion-answering1M<n<10M10 likes96 downloads3y agoHugging Face11OpenLab-NLP /tiny-instruct-kotextquestion-answering10K<n<100K1 likes93 downloads9mo agoHugging Face12Stur86 /tinyfacts Tinyfacts Short explanations of things, written using only about a thousand of the most common English words — the vocabulary Randall Munroe used for Thing Explainer, itself drawn from the xkcd comic Up Goer Five. Writing under that constraint forces a particular kind of prose. There is no word for photosynthesis, or gravity, or engine, so a text has to reach the idea by other means: green things that eat light, the way everything pulls on everything else, the part of the car… See the full description on the dataset page: https://huggingface.co/datasets/Stur86/tinyfacts.texttext-generation10K<n<100K0 likes70 downloads26d agoHugging Face13schneiderkamplab /dfm11-mathagentic-tinygsm-python dfm11-mathagentic-tinygsm-python English arithmetic word problems converted into native Python tool-call trajectories with precomputed tool responses and boxed final answers. Contents Rows: 367,749 Shards: 4 Format: deterministic gzip JSON Lines in data/train-*.jsonl.gz Schema: tools, four-message native tool trajectory, execution metadata, stable source ID, source revision, and admission status Intended repository:… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm11-mathagentic-tinygsm-python.texttext-generation100K<n<1M0 likes67 downloads23d agoHugging Face14psymon /Tiny-Ko-Stories Tiny-Ko-Stories English version is available below. Tiny-Ko-Stories는 TinyStories에서 영감을 받은 한국어 이야기 데이터셋입니다. TinyStories는 제한된 고품질 데이터셋을 사용하면, 소형 모델이라도 추론 능력과 창의력을 발휘할 수 있음을 보였습니다. 우리가 확인하려는 것은 단순합니다. 이 현상이 한국어에서도 재현될까? 이를 확인하려면 번역 데이터셋만으로는 부족했습니다. 한국어다운 이름, 문장 리듬, 의성어와 의태어, 색채어, 작은 사건 구조를 포함하려면 처음부터 한국어로 만든 이야기가 필요했습니다. 그래서 Tiny-Ko-Stories는 영어 TinyStories를 번역하는 대신, 한국어로 짧은 이야기를 새로 생성하고 여러 단계의 검수를 거쳐 구성했습니다. 데이터셋 요약 항목 값 레코드 수 2,003,542 형식 JSONL 공개… See the full description on the dataset page: https://huggingface.co/datasets/psymon/Tiny-Ko-Stories.texttext-generation1M<n<10M6 likes66 downloads4mo agoHugging Face15croqaz /tiny-vintage-completions Tiny vintage completions Synthetic vintage texts, with a cutoff date for year 1900. Based on unique 2-3 word seeds, extracted from croqaz/Vintage-v1, croqaz/Vintage-v2 and Haykgrigorian/English-historical-corpus-1800-1875. Check the files seeds1.txt and seeds2.txt. Generated by TypeWriter-7B-base and Talkie-13B-base completions. Citation If you find this dataset valuable, please consider citing: @misc{Tiny-vintage-completions, title = {Tiny vintage completions}… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/tiny-vintage-completions.tabulartext-generation100K<n<1M1 likes54 downloads22d agoHugging Face16jasonacox /tinychat_conversations TinyChat Conversations A specialized conversational dataset for midtraining language models with identity awareness and personality grounding. This dataset combines model identity conversations with question-answer pairs generated from personal blog content to infuse the model with a distinct conversational style and knowledge base. Dataset Description This dataset contains conversational data in JSONL format, designed for midtraining chat models alongside multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/jasonacox/tinychat_conversations.texttext-generation1K<n<10K1 likes53 downloads8mo agoHugging Face17tharunpranavsakthivel /TinyShell TinyShell Dataset TinyShell is a JSON Lines dataset of natural-language shell instructions paired with reference commands and semantic annotations. It covers Linux, macOS, and Windows examples across Bash, Zsh, and PowerShell-style environments. This release contains 65,848 records in three canonical splits: Split Records File Train 46,093 data/train.jsonl Validation 9,877 data/validation.jsonl Test 9,878 data/test.jsonl Release contents The… See the full description on the dataset page: https://huggingface.co/datasets/tharunpranavsakthivel/TinyShell.texttext-generation10K<n<100K0 likes42 downloads1mo agoHugging Face18recursal /SuperWiki-Tiny Dataset Card for SuperWiki-Tiny Waifu to catch your attention. Dataset Details Dataset Description SuperWiki-Tiny is a english only subset of the SuperWikipedia-NEXT (SuperWikiNEXT-32B) Curated by: KaraKaraWitch Funded by: Recursal.ai (I work there lol) Shared by: KaraKaraWitch Language(s) (NLP): English Only. License: cc-by-sa-4.0, Dataset Sources Source Data: https://dumps.wikimedia.org/other/enterprise_html/ Dataset Summary Refer… See the full description on the dataset page: https://huggingface.co/datasets/recursal/SuperWiki-Tiny.texttext-generation100K<n<1M0 likes41 downloads2y agoHugging Face19tinycomputerai /bun-server-bench-trajectories bun-server-bench trajectories Supervised fine-tuning and patch trajectories exported from bun-server-bench, a benchmark for evaluating coding agents on real-world Bun server engineering tasks. Every record comes from an agent run that passed both the public and hidden tests for its task — these are verified solutions, not raw attempts. The benchmark engineers each task so that a plausible-but-wrong implementation passes the visible tests and fails the hidden ones, so a passing… See the full description on the dataset page: https://huggingface.co/datasets/tinycomputerai/bun-server-bench-trajectories.texttext-generationn<1K1 likes37 downloads3mo agoHugging Face20llamafactory /xsum_tinyThis dataset is a subset of https://huggingface.co/datasets/EdinburghNLP/xsum. The training set is composed of 2,000 examples of the original training set and the test set is composed of 1,000 examples of the original validation set. textsummarization1K<n<10K0 likes36 downloads2y agoHugging Face21agentlans /starhopp3r-TinyChat starhopp3r/TinyChat This is an unofficial reformatted version of starhopp3r/TinyChat. It contains about 1 million conversations generated by GPT-4o mini using basic English. The conversations have been cleaned and put into ShareGPT-like format. Duplicates have been removed. All credit belongs to the original author. Footnote: These conversations have a strong mono no aware feeling in my opinion. texttext-generation100K<n<1M1 likes36 downloads11mo agoHugging Face22OpenLab-NLP /tiny-multiturn-chat-kotextquestion-answering1M<n<10M0 likes31 downloads10mo agoHugging Face23blueapple8259 /TinyFinewebEdu-kofineweb-edu에서 int_score가 4 이상인 데이터만 필터링한 후 DeepSeek-v3을 이용해 데이터를 단순한 형태의 문장과 간단한 어휘로 구성되도록 변환한 데이터입니다. 비용 문제로 데이터는 67k 개만 있습니다. texttext-generation10K<n<100K0 likes30 downloads2y agoHugging Face24MaxHastings /TinyLLMPretrainingCore Synthetic Simple-English Subject Explanations Dataset Dataset Summary This dataset contains synthetic, GPT-generated texts that explain a wide range of subjects using simple English.Each subject is expanded into multiple long-form explanations that repeat key ideas across different styles, perspectives, and framing strategies. The dataset is designed to emphasize clarity, redundancy, and consistency, making it useful for educational NLP, simplification tasks, and… See the full description on the dataset page: https://huggingface.co/datasets/MaxHastings/TinyLLMPretrainingCore.texttext-generation10K<n<100K0 likes25 downloads9mo agoHugging Face25Mattimax /TinyChat-ITA TinyChat-ITA Questo dataset fornisce coppie domanda-risposta in lingua italiana, pensate per applicazioni di chatbot e modelli conversazionali. Ogni entry contiene: Una domanda breve e naturale (input). Una risposta chiara e coerente (response). Tutti i dati sono memorizzati in formato JSONL, dove ogni riga rappresenta un oggetto JSON valido. Le risposte non parsabili sono state salvate separatamente per la revisione. Curato da: Mattimax per M.INC. (M.INC. profile) Condiviso… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/TinyChat-ITA.texttext-generation10K<n<100K0 likes23 downloads1y agoHugging Face26llamafactory /cnn_dailymail_tinyThis dataset is a subset of https://huggingface.co/datasets/cnn_dailymail. The training set is composed of 2,000 examples of the original training set and the test set is composed of 1,000 examples of the original validation set. We use the version 1.0.0 of the CNN/DailyMail dataset. textsummarization1K<n<10K0 likes22 downloads2y agoHugging Face27teleprint-me /tinypairs TinyPairs TinyPairs is a dataset of 1000 preprocessed input-target pairs derived from roneneldan/TinyStories.This dataset is formatted as a simple JSON file for easy use in training small-scale language models. 📜 Format: Each entry consists of: { "input": "Sue liked to study. She would sit upstairs in her room and look at her books.", "target": "Sue had a dog named Max. Max was deaf, but he was a good dog." } 🔧 How It Was Generated The dataset was extracted and… See the full description on the dataset page: https://huggingface.co/datasets/teleprint-me/tinypairs.texttext-generation10K<n<100K0 likes21 downloads2y agoHugging Face28purrgpt-community /The-Tiny-Purr-2 purrgpt-community/The-Tiny-Purr-2 The Tiny Purr is a turn base dataset that contains different lengths of conversation! What is cant do: Use tools Do web search (it can, hopefully) For what is: For a very frendly cat AI chatbot WARNING: There might be some mistakes texttext-generation1K<n<10K1 likes20 downloads1y agoHugging Face29llamafactory /adgen_tinyThis dataset is a subset of the advertising generation dataset proposed by https://aclanthology.org/D19-1321/. The training set is composed of 2,000 examples of the original training set and the test set is composed of 1,000 examples of the original validation set. texttext-generation1K<n<10K1 likes18 downloads2y agoHugging Face30V3N0M /TinyJenna-Uncensored-v01 Uncensored Alpaca Dataset: A New Frontier in Language Models This dataset is a collection of uncensored prompts and responses in the Alpaca format. It aims to provide a diverse and unfiltered source of data for training language models, pushing the boundaries of what these models can understand and generate. What Makes This Dataset Different? Uncensored: This dataset includes prompts and responses that touch upon topics that are often censored or avoided in traditional datasets.… See the full description on the dataset page: https://huggingface.co/datasets/V3N0M/TinyJenna-Uncensored-v01.texttext-generation1M<n<10M9 likes18 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.