CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bigcode /the-stack-smolgated Dataset Description A small subset (~0.1%) of the-stack dataset, each programming language has 10,000 random samples from the original dataset. The dataset has 2.6GB of text (code). Languages The dataset contains 30 programming languages: "assembly", "batchfile", "c++", "c", "c-sharp", "cmake", "css", "dockerfile", "fortran", "go", "haskell", "html", "java", "javascript", "julia", "lua", "makefile", "markdown", "perl", "php", "powershell", "python", "ruby", "rust"… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol.tabulartext-generation100K<n<1M95 likes27k downloads3y agoHugging Face02bigcode /the-stack-smol-xl Dataset Description A small subset of the-stack dataset, with 87 programming languages, each has 10,000 random samples from the original dataset. Languages The dataset contains 87 programming languages: 'ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c', 'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir', 'elm', 'emacs-lisp','erlang'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol-xl.tabulartext-generation100K<n<1M11 likes7.9k downloads4y agoHugging Face03bigcode /the-stack-smol-xs\tabulartext-generation1K<n<10K11 likes5.2k downloads4y agoHugging Face04opencsg /smoltalk-chinese Chinese SmolTalk Dataset [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report smoltalk-chinese is a Chinese fine-tuning dataset constructed with reference to the SmolTalk dataset. It aims to provide high-quality synthetic data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/smoltalk-chinese.tabulartext-generation10K<n<100K52 likes1.7k downloads10mo agoHugging Face05jordangong /jupyter-scripts-smollm3 The Stack v2 Jupyter Notebooks as Scripts This dataset contains script representations of the Jupyter notebooks in The Stack v2. It was created from the materialized Jupyter_Notebook split in jordangong/the-stack-v2-smollm3. The output schema follows the Jupyter-script schema used by bigcode/starcoderdata, but this release is not deduplicated, PII-filtered, or otherwise equivalent to StarCoderData's filtered split. Relationship to the SmolLM3 training mix This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.tabulartext-generation1M<n<10M0 likes1.1k downloads14d agoHugging Face06chankhavu /smolmo-sft-v2-seqlen64k smolmo-sft-v2-seqlen64k A supervised fine-tuning (SFT) dataset of math problems with full chain-of-thought solutions, formatted for the Olmo 3 "Thinking" models. 2,813,055 examples · ~37.9 B tokens. Three task families: proofs, numeric-answer problems, and tool-augmented (Python) problems. Every assistant turn carries an explicit <think> … </think> reasoning trace before the answer. Olmo 3 native chat + function-calling format; every example fits within a 64k-token context.… See the full description on the dataset page: https://huggingface.co/datasets/chankhavu/smolmo-sft-v2-seqlen64k.tabulartext-generation1M<n<10M0 likes924 downloads4mo agoHugging Face07AdaMLLab /smolkalam-arabic-conversational-sft SmolKalam SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets. Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.tabulartext-generation1M<n<10M3 likes890 downloads1mo agoHugging Face08bigcode /the-stack-v2-train-smol-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. bigcode/the-stack-v2-train-smol-ids: based on the bigcode/the-stack-v2-dedup… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids.tabulartext-generation10M<n<100M60 likes841 downloads2mo agoHugging Face09ginigen-ai /smol-worldcup 🏟️ Smol AI WorldCup — SHIFT Benchmark The world's first 5-axis evaluation framework for small language models. Not just "how smart?" — but "how honest? how fast? how small? how efficient?" 🏟️ Leaderboard huggingface.co/spaces/ginigen-ai/smol-worldcup 📊 Dataset huggingface.co/datasets/ginigen-ai/smol-worldcup 🏅 ALL Bench huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard 🏆 Official Ranking: WCS (WorldCup Score) WCS = √( SHIFT × PIR_norm )… See the full description on the dataset page: https://huggingface.co/datasets/ginigen-ai/smol-worldcup.tabulartext-generationn<1K47 likes420 downloads7mo agoHugging Face10sunorme /smoltalk-chinese Chinese SmolTalk Dataset [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report smoltalk-chinese is a Chinese fine-tuning dataset constructed with reference to the SmolTalk dataset. It aims to provide high-quality synthetic data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese… See the full description on the dataset page: https://huggingface.co/datasets/sunorme/smoltalk-chinese.tabulartext-generation10K<n<100K0 likes211 downloads6mo agoHugging Face11ChinaunicomSoftware /smoltalk-chinese-QwQ-Distrill smoltalk-chinese-QwQ-Distrill [中文] [English] 📖Technical Report smoltalk-chinese-QwQ-Distrill is a Chinese fine-tuning dataset constructed with reference to the SmolTalk-Chinese dataset. It aims to provide high-quality synthetic reasoning data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese LLMs across various tasks… See the full description on the dataset page: https://huggingface.co/datasets/ChinaunicomSoftware/smoltalk-chinese-QwQ-Distrill.tabulartext-generation100K<n<1M3 likes202 downloads2y agoHugging Face12neonforestmist /smolgpt-markdown-stories SmolGPT-Fables Stories A deterministic, text-only corpus of 96,000 original English Markdown stories built for SmolGPT-Fables. Every row is one complete supervised story example with an exact prompt / completion boundary, a requested scene count from one to six, and plain-language conditioning fields. No model, API, browser, or network service was used to create this dataset. Dataset summary 96,000 stories across 96,000 isolated story families 25 genres and all… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/smolgpt-markdown-stories.tabulartext-generation10K<n<100K0 likes136 downloads2mo agoHugging Face13nchapman /smoltalk-smol-magpie-ultra-no-refusals SmolTalk Smol-Magpie-Ultra No Refusals A Minos-cleaned version of HuggingFaceTB/smoltalk / smol-magpie-ultra for use as a neutral helpfulness SFT anchor. Rows are removed when NousResearch/Minos-v1 classifies the conversation as a refusal. The original train/test split structure is preserved. Cleaning version: minos-only-v1-2026-06-23 Counts Split Input rows Kept rows Dropped rows train 409,537 408,447 1,090 test 21,555 21,488 67 Overall removal… See the full description on the dataset page: https://huggingface.co/datasets/nchapman/smoltalk-smol-magpie-ultra-no-refusals.tabulartext-generation100K<n<1M1 likes114 downloads3mo agoHugging Face14BEE-spoke-data /the-stack-smol-xs-all bigcode/the-stack-smol-xs - all configs All configs from bigcode/the-stack-smol-xs concatenated and shuffled. 100 examples each of: ['ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c', 'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir', 'elm', 'emacs-lisp', 'erlang', 'f-sharp', 'fortran', 'glsl', 'go', 'groovy', 'haskell', 'html', 'idris', 'isabelle'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/the-stack-smol-xs-all.tabulartext-generation1K<n<10K0 likes56 downloads9mo agoHugging Face15pigandcat0624 /smoltalk-chinese Chinese SmolTalk Dataset [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report smoltalk-chinese is a Chinese fine-tuning dataset constructed with reference to the SmolTalk dataset. It aims to provide high-quality synthetic data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese… See the full description on the dataset page: https://huggingface.co/datasets/pigandcat0624/smoltalk-chinese.tabulartext-generation10K<n<100K0 likes38 downloads9mo agoHugging Face16pere /nb-asr-numerics-categorized-smoke-test Norwegian Bokmål Numeric Expression Categorized Dataset This dataset represents Stage 2 of the Norwegian numerics-data pipeline. It contains semantic validation and categorization annotations of the Norwegian numeric expression sentences harvested in Stage 1. Source Dataset Harvested Dataset: pere/nb-asr-numerics-harvested (approx. 3.8 million rows across 6 shards). Processing Architecture Inference Model: google/gemma-4-12B-it (instruction-tuned… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-categorized-smoke-test.tabulartext-generationn<1K0 likes24 downloads3mo agoHugging Face17Solshine /gemma-4-e2b-nla-eval-smoke Gemma-4-E2B NLA smoke-eval (20-row held-out set) A 20-row held-out subset of OpenWebText activations extracted from google/gemma-4-E2B at layer 23. Used as the canonical eval set for smoke-testing the v0.0.1 Gemma-4-E2B NLA pair on a fresh environment. This dataset is a subset of the held-out rl.parquet evaluation set used for the v0.0.1 round-trip eval (n=50 attempted, 42 evaluated after 8 empty-output exclusions, cos 0.438 ± 0.054). The 20-row subset preserves the activation… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-eval-smoke.tabulartext-generationn<1K0 likes22 downloads5mo agoHugging Face18SultanR /smolkalamgated Dataset Splits Gemma 3 Split Name # Examples LongAlign_64k_Qwen3_32B_yarn_131k_think 7,526 LongAlign_64k_context_lang_annotated_lang_6_no_think 6,249 Mixture_of_Thoughts_science_no_think 86,110 OpenHermes_2.5_no_think 384,900 OpenThoughts3_50K 50,000 OpenThoughts3_NoThink_180K 180,000 aya_dataset_Qwen3_32B_think 15,222 hermes_function_calling_v1_no_think 8,961 multi_turn_reasoning_if_think 28,217 s1k_1.1_think 835 smolagents_toolcalling_traces_think 9… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/smolkalam.tabulartext-generation10M<n<100M0 likes21 downloads10mo agoHugging Face19prometheus04 /matilda-smollm-mix-15b-gpt2 matilda-smollm-mix-15B-gpt2 15 B GPT-2-BPE tokens drawn from a 5:1 token-balanced mix of HuggingFaceTB/smollm-corpus: Source Share Tokens fineweb-edu-dedup 83.33 % 12.50 B cosmopedia-v2 16.67 % 2.50 B Total: 15,000,349,569 tokens across 151 shards (shard_*.bin, uint16, 100 M tokens per shard). The full SmolLM recipe is 75 / 15 / 10 fineweb-edu / cosmopedia-v2 / python-edu. python-edu was dropped because the HuggingFaceTB/smollm-corpus subset ships only blob_id… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/matilda-smollm-mix-15b-gpt2.tabulartext-generationn<1K1 likes19 downloads4mo agoHugging Face20syszxxxwu /smoltalk-chinese Chinese SmolTalk Dataset [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report smoltalk-chinese is a Chinese fine-tuning dataset constructed with reference to the SmolTalk dataset. It aims to provide high-quality synthetic data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese… See the full description on the dataset page: https://huggingface.co/datasets/syszxxxwu/smoltalk-chinese.tabulartext-generation10K<n<100K0 likes15 downloads5mo agoHugging Face21LocalDoc /smoltalk_azThis is part of the translated version of the original dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk tabulartranslation100K<n<1M1 likes13 downloads10mo agoHugging Face22habibahabchi /smollm3-base-blindspots SmolLM3-3B-Base Blind Spots Evaluation Dataset Dataset Summary This dataset documents 10 diverse failure cases discovered while evaluating HuggingFaceTB/SmolLM3-3B-Base, a 3-billion parameter decoder-only base language model released by Hugging Face in July 2025. The evaluation was conducted as part of the Fatima Fellowship technical challenge on Blind Spots of Frontier Models. Model Tested Model: HuggingFaceTB/SmolLM3-3B-Base Parameters: 3 billion… See the full description on the dataset page: https://huggingface.co/datasets/habibahabchi/smollm3-base-blindspots.tabulartext-generationn<1K0 likes8 downloads6mo agoHugging Face23BEE-spoke-data /smollm-corpus-pythongated smollm-corpus - python A version of the python-edu subset with the text added tabulartext-generation10M<n<100M0 likes6 downloads9mo agoHugging Face2413point5 /reverse-text-tinystories-easy-smoke Reverse Text TinyStories Easy Smoke This is a small smoke-test dataset for the reverse-text task. Splits train: 12 rows test: 4 rows Columns prompt char_count word_count source Source Derived from roneneldan/TinyStories by taking non-overlapping word windows and keeping only prompts that fall in the easy character-length bucket. Difficulty Rule All rows in this dataset are easy examples with prompt lengths in the 20-74 character range.… See the full description on the dataset page: https://huggingface.co/datasets/13point5/reverse-text-tinystories-easy-smoke.tabulartext-generationn<1K0 likes5 downloads6mo agoHugging Face25FineEnvs /SmolDataEnvs-sft 🛠️ SmolDataEnvs — SFT 5.5K+ RL tasks for hill-climbing small models in code and data science. A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on. Two runs over the same 5,000 tasks — shuffled against a curriculum ordered easiest to hardest. 4,677 worked examples of an agent doing data science the right way. Each row is a complete, verified-correct trajectory: read the question, poke at the data with a shell tool… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/SmolDataEnvs-sft.tabulartext-generation1K<n<10K0 likes10m agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.