CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bigcode /the-stack-smolgated Dataset Description A small subset (~0.1%) of the-stack dataset, each programming language has 10,000 random samples from the original dataset. The dataset has 2.6GB of text (code). Languages The dataset contains 30 programming languages: "assembly", "batchfile", "c++", "c", "c-sharp", "cmake", "css", "dockerfile", "fortran", "go", "haskell", "html", "java", "javascript", "julia", "lua", "makefile", "markdown", "perl", "php", "powershell", "python", "ruby", "rust"… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol.tabulartext-generation100K<n<1M95 likes27k downloads3y agoHugging Face02Avelina /smollm-corpus-cleaned SmolLM-Corpus: Now shuffled and sharded (and Cleaned)! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.texttext-generation100M<n<1B2 likes11k downloads2y agoHugging Face03bigcode /the-stack-smol-xl Dataset Description A small subset of the-stack dataset, with 87 programming languages, each has 10,000 random samples from the original dataset. Languages The dataset contains 87 programming languages: 'ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c', 'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir', 'elm', 'emacs-lisp','erlang'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol-xl.tabulartext-generation100K<n<1M11 likes7.9k downloads4y agoHugging Face04ginigen-ai /smol-worldcup 🏟️ Smol AI WorldCup — SHIFT Benchmark The world's first 5-axis evaluation framework for small language models. Not just "how smart?" — but "how honest? how fast? how small? how efficient?" 🏟️ Leaderboard huggingface.co/spaces/ginigen-ai/smol-worldcup 📊 Dataset huggingface.co/datasets/ginigen-ai/smol-worldcup 🏅 ALL Bench huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard 🏆 Official Ranking: WCS (WorldCup Score) WCS = √( SHIFT × PIR_norm )… See the full description on the dataset page: https://huggingface.co/datasets/ginigen-ai/smol-worldcup.tabulartext-generationn<1K47 likes420 downloads7mo agoHugging Face05enPurified /smollm-corpus-fineweb-edu-enPurified-openai-messages 📖 smollm-corpus-fineweb-edu-enPurified-openai-messages smollm-corpus-fineweb-edu-enPurified is a highly curated, "prose-first" subset of the fineweb-edu-dedup subset found in HuggingFaceTB/smollm-corpus. The enPurified collection is built on a specific philosophy: Specialization. While the original dataset is excellent for general pre-training, high-quality fluent English prose often gets diluted when mixed with syntax-heavy code, rigid math formulas, or low-information web junk.… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-fineweb-edu-enPurified-openai-messages.texttext-generation10M<n<100M1 likes338 downloads8mo agoHugging Face06Makoveli89 /bigcode-the-stack-smol bigcode/the-stack-smol This repository documents a dataset used by the Mothership project. By default data is not mirrored here. Primary source: https://huggingface.co/datasets/bigcode/the-stack-smol Local cache (if present during publishing): C:\Users\Sean Smith\Documents\Scraps\Knowledge\Mothership\library\datasets\bigcode\the-stack-smol Revision pin: none To reproduce locally, use the project's downloader: python scripts/download_datasets.py --include bigcode/the-stack-smol textn<1K0 likes285 downloads1y agoHugging Face07Cornerf /rebot-can-sort-stage1-v1-smoke ReBot can sorting Stage 1 Reviewed success-only LeRobot v3 dataset for: Pick up one can and place it in the taped sorting zone. Episodes: 52 Frames: 36729 FPS: 30 Robot: seeed_b601_dm_follower Cameras: observation.images.front (Logitech overhead) and observation.images.side (Innomaker wrist/claw) Action order: shoulder_pan.pos, shoulder_lift.pos, elbow_flex.pos, wrist_flex.pos, wrist_yaw.pos, wrist_roll.pos, gripper.pos Intended destination:… See the full description on the dataset page: https://huggingface.co/datasets/Cornerf/rebot-can-sort-stage1-v1-smoke.tabularroboticsn<1K0 likes258 downloads2mo agoHugging Face08enPurified /smollm-corpus-cosmopedia-v2-enPurified-openai-messages enPurified Collection: Smollm Corpus Cosmopedia V2] Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows. Purpose of the enPurified Collection The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.texttext-generation1M<n<10M1 likes230 downloads8mo agoHugging Face09ChinaunicomSoftware /smoltalk-chinese-QwQ-Distrill smoltalk-chinese-QwQ-Distrill [中文] [English] 📖Technical Report smoltalk-chinese-QwQ-Distrill is a Chinese fine-tuning dataset constructed with reference to the SmolTalk-Chinese dataset. It aims to provide high-quality synthetic reasoning data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese LLMs across various tasks… See the full description on the dataset page: https://huggingface.co/datasets/ChinaunicomSoftware/smoltalk-chinese-QwQ-Distrill.tabulartext-generation100K<n<1M3 likes202 downloads2y agoHugging Face10Cornerf /rebot-two-can-recycle-v2-smoke ReBot can sorting Stage 1 Reviewed success-only LeRobot v3 dataset for: Pick up one can and place it in the taped sorting zone. Episodes: 25 Frames: 14478 FPS: 30 Robot: seeed_b601_dm_follower Cameras: observation.images.front (Logitech overhead) and observation.images.side (Innomaker wrist/claw) Action order: shoulder_pan.pos, shoulder_lift.pos, elbow_flex.pos, wrist_flex.pos, wrist_yaw.pos, wrist_roll.pos, gripper.pos Intended destination:… See the full description on the dataset page: https://huggingface.co/datasets/Cornerf/rebot-two-can-recycle-v2-smoke.tabularroboticsn<1K0 likes179 downloads2mo agoHugging Face11zwanderer /smolvlm2-fire-videostext1K<n<10K0 likes128 downloads6mo agoHugging Face12enPurified /smoltalk-creative-writing-enPurified-openai-messages 📖 SmolTalk-Creative-Writing-enPurified-openai-messages SmolTalk-Creative-Writing-enPurified is a highly curated, "prose-first" subset of the original collinear-ai/smoltalk-creative-writing dataset. The enPurified collection is built on a specific philosophy: Specialization. While the ecosystem has plenty of datasets for coding (StackOverflow, StarCoder) and mathematics (GSM8K), high-quality, fluent English prose often gets diluted when mixed with syntax-heavy code or rigid math… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smoltalk-creative-writing-enPurified-openai-messages.texttext-generation10K<n<100K2 likes115 downloads9mo agoHugging Face13dougalldeepmind /2026-09-10-delib-synth-smoke Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7) field value experiment Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7) date_generated 20260910_191448 constitution constitutions/abridged/constitution.md source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-10-delib-synth-smoke.tabularn<1K0 likes105 downloads14d agoHugging Face14bestdive /details_bestdive__SmolLM3-3B-SFT-Free-Course Smol course SFT evaluation - Kay Zheng Actual full GSM8K test evaluation of bestdive/SmolLM3-3B-SFT-Free-Course, adapter revision 0484e028b494d605a267050a949c9266edadd16b, merged with pinned SmolLM3-3B-Base before evaluation. Full 1319 test examples, zero-shot, original extractive_match: 0.4086429112964367 (stderr 0.013540639733342422). Free Google Colab T4, no paid HF Jobs; cost 0. lighteval 0.11.0, vLLM 0.10.1.1, Transformers 4.57.1, Python 3.12. Dataset-address correction… See the full description on the dataset page: https://huggingface.co/datasets/bestdive/details_bestdive__SmolLM3-3B-SFT-Free-Course.textn<1K0 likes100 downloads14d agoHugging Face15dongboklee /test_smollm MMLU-Pro Multi-Domain Dataset: test_smollm Usage from datasets import load_dataset # Load entire dataset dataset = load_dataset("dongboklee/test_smollm") # Load specific domain law_dataset = load_dataset("dongboklee/test_smollm", split="law") text1K<n<10K0 likes92 downloads1y agoHugging Face16dougalldeepmind /2026-09-11-delib-sonnet-synth-smoke Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7) field value experiment Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7) date_generated 20260911_192737 constitution constitutions/abridged/constitution.md source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-11-delib-sonnet-synth-smoke.tabularn<1K0 likes81 downloads13d agoHugging Face17Nebulous /lmsys-chat-1m-smortmodelsonlyThis version of the dataset only has responses from GPT-4, Claude-1, Claude-2, Claude-instant-1, and GPT-3.5-turbo text10K<n<100K4 likes72 downloads3y agoHugging Face18zjxia /perfectblend-smoltalk-chinese-large-blend-regentext100K<n<1M0 likes71 downloads6mo agoHugging Face19dougalldeepmind /2026-09-17-da-lowstakes-constitution-synth-smoke 18-row constitution-only low-stakes smoke; FAILED scaling gate; diagnostic candidates only field value experiment 18-row constitution-only low-stakes smoke; FAILED scaling gate; diagnostic candidates only date_generated 20260917_145731 constitution constitutions/claude_distilled_09_principles/constitution.md sha256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc source_repo https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-da-lowstakes-constitution-synth-smoke.textn<1K0 likes71 downloads7d agoHugging Face20dougalldeepmind /2026-09-17-da-lowstakes-values-in-advice-synth-smoke 18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only field value experiment 18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only date_generated 20260917_171453 constitution constitutions/claude_distilled_09_principles/constitution.md sha256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc source_repo https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-da-lowstakes-values-in-advice-synth-smoke.textn<1K0 likes70 downloads7d agoHugging Face21dougalldeepmind /2026-09-09-delib-synth-smoke Deliberative SFT: delib; native Qwen reasoning, no judge filtering field value experiment Deliberative SFT: delib; native Qwen reasoning, no judge filtering date_generated 20260909_152646 constitution constitutions/abridged_no_delib/constitution.md source_repo https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 5302fa9589d32c4296d9393b4838a00d91042b76 models qwen/qwen3.6-27b through Alibaba/OpenRouter (API revision not exposed)… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-09-delib-synth-smoke.textn<1K0 likes68 downloads15d agoHugging Face22malaiwah /qfs-smollm2-135m-wikitext2-native-v1 HF workflow d3dc69602aeb981f06bd9f4c726937f9 A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-native-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-native-v1.tabularn<1K0 likes67 downloads16d agoHugging Face23dougalldeepmind /2026-09-17-delib-noref-synth-smoke Deliberative SFT: delib-noref; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7) field value experiment Deliberative SFT: delib-noref; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7) date_generated 20260917_191247 constitution constitutions/abridged/constitution.md source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-delib-noref-synth-smoke.tabularn<1K0 likes67 downloads7d agoHugging Face24ychen /Generated-Empathetic-Dialogues-v0.1-Smol Generated Empathetic Conversations v0.1 - Smol This is dataset contains 10K rows of multi-round empathetic conversations convering a diverse set of topics. Highlights Multi-round conversation It's not single-turn. The user and the assistant works together to gradually unfold the conversation. The average number of turns is 5, with a standard deviation of approximately 1.59 turns. A turn consists of two messages with one by the user, and another by the… See the full description on the dataset page: https://huggingface.co/datasets/ychen/Generated-Empathetic-Dialogues-v0.1-Smol.texttext-generation10K<n<100K4 likes66 downloads2y agoHugging Face25costadev00 /smoke-openai-terra-batch-brasil-25-20260724-01 Smoke OpenAI Terra Batch — Brasil × 25 tasks Run real de validação do fluxo matricial document_task_matrix, executada sobre um único documento da Wikipédia em português com o título Brasil. Cada uma das 25 tasks canônicas recebeu exatamente um slot inicial. Resultado status: completed documentos: 1 pares planejados: 25 exemplos aceitos: 25 pares pulados: 0 pares esgotados: 0 resultados reais do backend: 27 retries com nova chamada: 2 backend: openai_api… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/smoke-openai-terra-batch-brasil-25-20260724-01.texttext-generationn<1K0 likes61 downloads2mo agoHugging Face26open-llm-leaderboard /HuggingFaceTB__SmolLM-1.7B-Instruct-detailsgated Dataset Card for Evaluation run of HuggingFaceTB/SmolLM-1.7B-Instruct Dataset automatically created during the evaluation run of model HuggingFaceTB/SmolLM-1.7B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceTB__SmolLM-1.7B-Instruct-details.tabular10K<n<100K0 likes59 downloads2y agoHugging Face27dougalldeepmind /2026-08-04-ddp-smoke-bundleDDP smoke bundle: train_lora.py + a 64-example toy set, for validating multi-GPU wiring. textn<1K0 likes57 downloads2mo agoHugging Face28dougalldeepmind /2026-09-08-delib-synth-smoke Deliberative SFT: delib; native Qwen reasoning, no judge filtering field value experiment Deliberative SFT: delib; native Qwen reasoning, no judge filtering date_generated 20260908_182742 constitution constitutions/no_claude_mentioned/constitution.md source_repo https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ a25750eb516cb6254640ae39a3a870f49398dbf0 models qwen/qwen3.6-27b through Alibaba/OpenRouter (API revision not exposed)… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-08-delib-synth-smoke.textn<1K0 likes57 downloads16d agoHugging Face29someone13574 /smoltalk-binidxtext1M<n<10M1 likes56 downloads2y agoHugging Face30sbussiso /SmolThinker-Synthetic-Reasoning SmolThinker Synthetic Reasoning Dataset 7,796 examples that teach a small instruct model to emit its reasoning inside literal <think> ... </think> blocks, so chat UIs that render collapsible reasoning (Open WebUI, Ollama, LM Studio) pick them up. Model-agnostic: nothing in the data names a specific model. Reasoning length scales with task difficulty, from one line on a greeting to 2,000+ characters on a hard question. from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/SmolThinker-Synthetic-Reasoning.texttext-generation10K<n<100K1 likes55 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.