CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01roneneldan /TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/roneneldan/TinyStories.texttext-generation1M<n<10M1.2k likes95k downloads2y agoHugging Face02llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes45k downloads2y agoHugging Face03Trelis /tiny-shakespeare Data source Downloaded via Andrej Karpathy's nanogpt repo from this link Data Format The entire dataset is split into train (90%) and test (10%). All rows are at most 1024 tokens, using the Llama 2 tokenizer. All rows are split cleanly so that sentences are whole and unbroken. texttext-generationn<1K11 likes14k downloads3y agoHugging Face04ysn-rfd /text-dataset-tiny-code-script-py-format USED of tahamajs/medicine_ds_persian for .parquet file USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file USED of Abirate/english_quotes for .jsonl file NEW FILES (05/12/2025) NEW FILES (12/26/2025) NEW FILES (02/15/2026) texttext-generation10K<n<100K3 likes1.8k downloads4mo agoHugging Face05nampdn-ai /tiny-codesgated Reasoning with Language and Code This synthetic dataset is a collection of 1.6 millions short and clear code snippets that can help LLM models learn how to reason with both natural and programming languages. The dataset covers a wide range of programming languages, such as Python, TypeScript, JavaScript, Ruby, Julia, Rust, C++, Bash, Java, C#, and Go. It also includes two database languages: Cypher (for graph databases) and SQL (for relational databases) in order to study the… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-codes.texttext-generation1M<n<10M302 likes1.6k downloads3y agoHugging Face06CohereLabs /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Data source Prompts from AM-DeepSeek-R1-0528-Distilled Thinking traces and outputs distilled from gpt-oss-120b Translated with command-a-translate and DeepSeek-V3 Languages (44) Language Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M6 likes1.2k downloads16d agoHugging Face07nampdn-ai /tiny-textbooksgated Textbook-like Dataset: A High-Quality Resource for Small Language Models The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model. Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.tabulartext-generation100K<n<1M184 likes567 downloads2y agoHugging Face08fzmnm /TinyStoriesAdv-zh TinyStoriesAdv keywords: grade school level, large language model, small language model, tiny language model, super tiny language model, 小学生知识水平,大语言模型,小语言模型,迷你语言模型, llm, slm. 受到TinyStories、Phi2等论文的启发,我制作了一个约1B tokens的小学知识水平的“一揽子”大语言模型训练语料库。 “一揽子”指的是本数据集是众多数据集的集合。为了提升模型的不同能力(例如事实性知识、元认知、思维链、阅读理解RAG、逻辑推理等),我开了不少脑洞,使用了多种创新的提示词生成了具有多样性和针对性的子数据集。… See the full description on the dataset page: https://huggingface.co/datasets/fzmnm/TinyStoriesAdv-zh.texttext-generation100M<n<1B11 likes488 downloads2y agoHugging Face09fhswf /TinyStoriesV2_cleaned License: CDLA-Sharing-1.0 Dataset containing synthetically generated (GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. This is a cleaned up Version of the original TinyStories Dataset: https://huggingface.co/datasets/roneneldan/TinyStories. We thank the authors for their contribution. This Version only contains cleaned-up stories generated by GPT4. Stories were deleted that contained spelling and… See the full description on the dataset page: https://huggingface.co/datasets/fhswf/TinyStoriesV2_cleaned.texttext-generation1M<n<10M13 likes398 downloads2y agoHugging Face10AL-GR /AL-GR-Tiny AL-GR-Tiny: A Complete & Sampled Generative Recommendation Dataset Dataset Summary AL-GR-Tiny is a compact, self-contained, and sampled version of the large-scale AL-GR ecosystem. It is designed for users who want to quickly experiment, develop, or understand the full pipeline of generative recommendation without needing to process terabytes of data. This "all-in-one" repository bundles everything you need: Pre-processed Training/Testing Data: Ready-to-use data for… See the full description on the dataset page: https://huggingface.co/datasets/AL-GR/AL-GR-Tiny.texttext-generation10M<n<100M3 likes397 downloads1y agoHugging Face11fzmnm /TinyHelen-zh What's New Mar.31 2025 Added instruct fine-tuning and reasoning dataset in the same ELI5 style. Take a look! Mar.31 2025 See my new model. TinyHelen-zh Inspired by the paper TinyHelen's First Curriculum, we present a Chinese version of the LLM-simplified training corpus. This dataset is converted from high-quality Chinese and English web crawls for training baby-size (<100M) language models. Adult-talking 北京市财政局、北京海关、国家税务总局北京市税务局、北京市国际服务贸易事务中心:… See the full description on the dataset page: https://huggingface.co/datasets/fzmnm/TinyHelen-zh.texttext-generation100K<n<1M1 likes390 downloads1y agoHugging Face12VatsaDev /TinyTextThe entire NanoPhi Dataset is at train.jsonl Separate Tasks Include Math (Metamath, mammoth) Code (Code Search Net) Logic (Open-platypus) Roleplay (PIPPA, RoleplayIO) Textbooks (Tiny-text, Sciphi) Textbook QA (Orca-text, Tiny-webtext) textquestion-answering1M<n<10M34 likes364 downloads2y agoHugging Face13Gabrui /multilingual_TinyStories Dataset Card for Multilingual TinyStories Dataset Details Dataset Description The Multilingual TinyStories dataset contains translations of the original TinyStories dataset, which consists of synthetically generated short stories using a small vocabulary suitable for 3 to 4-year-olds. These stories were originally generated by GPT-3.5 and GPT-4. The multilingual versions have been translated into various languages, including Spanish, Chinese, German, Turkish… See the full description on the dataset page: https://huggingface.co/datasets/Gabrui/multilingual_TinyStories.texttext-generation10M<n<100M1 likes345 downloads2y agoHugging Face14Dxniz /TinyStories-Multilingual Novelist: TinyStories Multilingual Edition Dataset Summary The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short, child-safe fiction designed to stress-test literary consistency, emotional warmth, and multilingual fluency in small models. Derived from the broader Novelist ecosystem, this subset focuses on narrative simplicity paired with complex moral and social themes. The dataset contains 15,688 high-quality stories across 28 languages. Each… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/TinyStories-Multilingual.texttext-generation10K<n<100K1 likes334 downloads6mo agoHugging Face15erenyeager-1 /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Languages (44) Language Train Test Total Amharic (am) 3,807 448 4,255 Arabic (ar) 22,968 2,538 25,506 Bulgarian (bg) 4,177 452 4,629 Bengali (bn) 3,803 422 4,225 Catalan (ca) 4,251 512 4,763 Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M0 likes323 downloads17d agoHugging Face16fzmnm /TinyBooks-QA-Chinese本数据集已停止更新,请移步https://huggingface.co/datasets/fzmnm/TinyStoriesAdv-zh TinyBooks-QA-Chinese Inspired by the (TinyStories)[https://arxiv.org/abs/2305.07759] paper, where a small language model exhibits strong capabilities when trained on high-quality, 🍼baby-friendly stories synthesized by AI, I present an AI-generated Encyclopedia suitable for kindergarten and grade school levels. This AI-synthesized dataset converts classical literature into a question-answer style curriculum with… See the full description on the dataset page: https://huggingface.co/datasets/fzmnm/TinyBooks-QA-Chinese.texttext-generation1K<n<10K8 likes317 downloads2y agoHugging Face17SauravP97 /tiny-stories-tokenized-bpetexttext-generation1M<n<10M1 likes299 downloads7mo agoHugging Face18taesiri /TinyStories-Farsi Tiny Stories Farsi The Tiny Stories Farsi project is a continuous effort to translate the Tiny Stories dataset into the Persian (Farsi) language. The primary goal is to produce a high-quality Farsi dataset, maintaining equivalency with the original English version, and subsequently to utilize it for training language models in Farsi. This seeks to affirm that the advancements and trends observed in English language models are replicable and applicable in other languages. Thus far… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/TinyStories-Farsi.texttext-generation100K<n<1M18 likes272 downloads3y agoHugging Face19malaiwah /k2-horizon-tiny-cpu-repro-v1 K2-Horizon MoVA tiny random CPU fixture Complete untrained K2HorizonForCausalLM, not Moonshot Kimi despite the K2 name. Architecture source: IFM/K2-Horizon-MoVA-36B-A4B at 05cab0a4d7150c1c460a000b37ff40cc1af2feaa. No pretrained weights, original tokenizer, training data, paid GPU or cloud compute used. Complete text-only K2HorizonForCausalLM, not Kimi: three-layer dense prefix followed by two real MoVA+MoE layers, grouped RMSNorm, sigmoid top-k routing with selection-only bias… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/k2-horizon-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes271 downloads19d agoHugging Face20malaiwah /deepseek-v4-tiny-cpu-repro-v1 DeepSeek-V4 tiny corrected-native-primitives CPU text fixture Complete randomly initialized, untrained QFSDeepseekV4ForCausalLM text class using Transformers5.16.1 native primitives and a reviewed RMSNorm arithmetic correction. No upstream weights, paid GPU/cloud compute or useful-model claim. This is not unmodified native Transformers or the complete production release. Architecture and scope Text lineage:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/deepseek-v4-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes270 downloads19d agoHugging Face21HayatoHongo /TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/HayatoHongo/TinyStories.texttext-generation1M<n<10M0 likes265 downloads10mo agoHugging Face22malaiwah /qwen3-5-tiny-cpu-repro-v1 Qwen3.5 tiny native random CPU fixture Complete randomly initialized, untrained Qwen3_5ForConditionalGeneration checkpoint. This is a pipeline/reproducibility fixture, not a useful language model, distillation, quantization, quality benchmark, or claim about the performance of Qwen3.8-27B. No upstream model weights or training data were used. No paid GPU/cloud compute. Architecture and lineage Architecture lineage: Qwen/Qwen3.8-27B at… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-5-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes251 downloads19d agoHugging Face23malaiwah /glm5-next-tiny-cpu-repro-v1This repository is an evidence bundle, not one root-format dataset at repository root. first/ and repeat/ are separate complete sealed QFS root datasets; comparison/ holds the comparison receipt and tokenwise result. panel/ is the sealed input panel. Other files are provenance, logs and reproduction tools. Do not pass the bundle root as a QFS dataset. GLM5-Next tiny native CPU fixture This is a complete untrained random-initialized native Glm5NextForConditionalGeneration wrapper… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm5-next-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes243 downloads19d agoHugging Face24ivanzhouyq /RedPajama-Tiny Dataset Card for Dataset Name Dataset Summary This is a tiny version of the RedPajama dataset. It contains 64 samples from each of the 7 sources. This dataset is intended for developing and testing data/training pipeline for loading the full RedPajama dataset or any general HuggingFace dataset. It is very fast to download and easy to examine. You should not use it for training a full model, but you can use it for overfitting test or any other sanity checks.… See the full description on the dataset page: https://huggingface.co/datasets/ivanzhouyq/RedPajama-Tiny.texttext-generationn<1K5 likes216 downloads2y agoHugging Face25karmx /TinyQuery-Tools-Multilingual TinyQuery Tools Multilingual GitHub: source code, setup guide, streaming examples and tests A synthetic, fictitious dataset for learning schema-conditioned SQL and tool actions from English, imperfect English, Hindi and Hinglish. Created for a four-hour, from-scratch small-model experiment. It contains no real user databases. Split Examples Purpose Train 347,376 Semantic scenarios, teacher language and randomized context variants Validation 1,200 Held-out domain… See the full description on the dataset page: https://huggingface.co/datasets/karmx/TinyQuery-Tools-Multilingual.texttext-generation100K<n<1M1 likes210 downloads13d agoHugging Face26styfeng /TinyDialogues Dataset Card for TinyDialogues TinyDialogues dataset collected as part of the EMNLP 2024 paper "Is Child-Directed Speech Effective Training Data for Language Models?" by Steven Y. Feng, Noah D. Goodman, and Michael C. Frank. For more details, please see Appendices A-C in our paper. Dataset Sources Repository: https://github.com/styfeng/TinyDialogues Paper: https://aclanthology.org/2024.emnlp-main.1231/ Dataset Structure Final training and validation data… See the full description on the dataset page: https://huggingface.co/datasets/styfeng/TinyDialogues.texttext-generation100K<n<1M1 likes205 downloads2y agoHugging Face27AlexKitipov /TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/AlexKitipov/TinyStories.texttext-generation1M<n<10M0 likes202 downloads4mo agoHugging Face28exnivo /tinybrain-instruct-sft-200k TinyBrain Instruct 200K A 196k+ row English SFT dataset for training tiny instruction-following language models. TinyBrain Instruct 200K is a synthetic supervised fine-tuning dataset made for small language models, especially models around 100M–500M parameters. The dataset focuses on short, clear, learnable assistant responses across education, basic math reasoning, clean conversation, planning, simplification, simple coding, and honesty/uncertainty behavior. Most… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/tinybrain-instruct-sft-200k.texttext-generation100K<n<1M3 likes178 downloads3mo agoHugging Face29nampdn-ai /tiny-strange-textbooksgated Quirky Textbook Trove: Compact Excellence for Small Language Model Strange dataset is 100% AI-generated, a compilation aligned with the vision of the Textbooks Are All You Need and Textbooks Are All You Need II: phi-1.5 technical report research. This dataset features 2,7M synthetic textbooks, encapsulating 16GB of raw text data. The unique name reflects its unconventional synthesis methodology, its compact size, deduped, and its emphasis on clear, focused content. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-strange-textbooks.texttext-generation1M<n<10M94 likes176 downloads3y agoHugging Face300rn0 /tinystories-instruct-balanced Dataset Card for TinyStories Instruct - Balanced Dataset Summary TinyStories Instruct - Balanced is a curated, instruction-tuning dataset derived from roneneldan/TinyStoriesInstruct. It contains short story generation examples with balanced happy/sad endings (50-50 split), making it ideal for fine-tuning language models to follow instructions and generate contextually appropriate narratives. The dataset was created to address the original TinyStoriesInstruct's… See the full description on the dataset page: https://huggingface.co/datasets/0rn0/tinystories-instruct-balanced.texttext-generation100K<n<1M0 likes176 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.