datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-stack-smol
Dataset Description
A small subset (~0.1%) of the-stack dataset, each programming language has 10,000 random samples from the original dataset. The dataset has 2.6GB of text (code).
Languages
The dataset contains 30 programming languages:
"assembly", "batchfile", "c++", "c", "c-sharp", "cmake", "css", "dockerfile", "fortran", "go", "haskell", "html", "java",
"javascript", "julia", "lua", "makefile", "markdown", "perl", "php", "powershell", "python", "ruby", "rust"… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol.smollm-corpus-cleaned
SmolLM-Corpus: Now shuffled and sharded (and Cleaned)!
This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming!
The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo.
Dataset Structure
The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.the-stack-smol-xl
Dataset Description
A small subset of the-stack dataset, with 87 programming languages, each has 10,000 random samples from the original dataset.
Languages
The dataset contains 87 programming languages:
'ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c',
'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir',
'elm', 'emacs-lisp','erlang'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol-xl.smol-worldcup
🏟️ Smol AI WorldCup — SHIFT Benchmark
The world's first 5-axis evaluation framework for small language models.
Not just "how smart?" — but "how honest? how fast? how small? how efficient?"
🏟️ Leaderboard
huggingface.co/spaces/ginigen-ai/smol-worldcup
📊 Dataset
huggingface.co/datasets/ginigen-ai/smol-worldcup
🏅 ALL Bench
huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard
🏆 Official Ranking: WCS (WorldCup Score)
WCS = √( SHIFT × PIR_norm )… See the full description on the dataset page: https://huggingface.co/datasets/ginigen-ai/smol-worldcup.smollm-corpus-fineweb-edu-enPurified-openai-messages
📖 smollm-corpus-fineweb-edu-enPurified-openai-messages
smollm-corpus-fineweb-edu-enPurified is a highly curated, "prose-first" subset of the fineweb-edu-dedup subset found in HuggingFaceTB/smollm-corpus.
The enPurified collection is built on a specific philosophy: Specialization. While the original dataset is excellent for general pre-training, high-quality fluent English prose often gets diluted when mixed with syntax-heavy code, rigid math formulas, or low-information web junk.… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-fineweb-edu-enPurified-openai-messages.bigcode-the-stack-smol
bigcode/the-stack-smol
This repository documents a dataset used by the Mothership project. By default data is not mirrored here.
Primary source: https://huggingface.co/datasets/bigcode/the-stack-smol
Local cache (if present during publishing): C:\Users\Sean Smith\Documents\Scraps\Knowledge\Mothership\library\datasets\bigcode\the-stack-smol
Revision pin: none
To reproduce locally, use the project's downloader:
python scripts/download_datasets.py --include bigcode/the-stack-smol
rebot-can-sort-stage1-v1-smoke
ReBot can sorting Stage 1
Reviewed success-only LeRobot v3 dataset for: Pick up one can and place it in
the taped sorting zone.
Episodes: 52
Frames: 36729
FPS: 30
Robot: seeed_b601_dm_follower
Cameras: observation.images.front (Logitech overhead) and
observation.images.side (Innomaker wrist/claw)
Action order: shoulder_pan.pos, shoulder_lift.pos, elbow_flex.pos, wrist_flex.pos, wrist_yaw.pos, wrist_roll.pos, gripper.pos
Intended destination:… See the full description on the dataset page: https://huggingface.co/datasets/Cornerf/rebot-can-sort-stage1-v1-smoke.smollm-corpus-cosmopedia-v2-enPurified-openai-messages
enPurified Collection: Smollm Corpus Cosmopedia V2]
Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows.
Purpose of the enPurified Collection
The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.smoltalk-chinese-QwQ-Distrill
smoltalk-chinese-QwQ-Distrill [中文] [English]
📖Technical Report
smoltalk-chinese-QwQ-Distrill is a Chinese fine-tuning dataset constructed with reference to the SmolTalk-Chinese dataset. It aims to provide high-quality synthetic reasoning data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese LLMs across various tasks… See the full description on the dataset page: https://huggingface.co/datasets/ChinaunicomSoftware/smoltalk-chinese-QwQ-Distrill.rebot-two-can-recycle-v2-smoke
ReBot can sorting Stage 1
Reviewed success-only LeRobot v3 dataset for: Pick up one can and place it in
the taped sorting zone.
Episodes: 25
Frames: 14478
FPS: 30
Robot: seeed_b601_dm_follower
Cameras: observation.images.front (Logitech overhead) and
observation.images.side (Innomaker wrist/claw)
Action order: shoulder_pan.pos, shoulder_lift.pos, elbow_flex.pos, wrist_flex.pos, wrist_yaw.pos, wrist_roll.pos, gripper.pos
Intended destination:… See the full description on the dataset page: https://huggingface.co/datasets/Cornerf/rebot-two-can-recycle-v2-smoke.smolvlm2-fire-videossmoltalk-creative-writing-enPurified-openai-messages
📖 SmolTalk-Creative-Writing-enPurified-openai-messages
SmolTalk-Creative-Writing-enPurified is a highly curated, "prose-first" subset of the original collinear-ai/smoltalk-creative-writing dataset.
The enPurified collection is built on a specific philosophy: Specialization. While the ecosystem has plenty of datasets for coding (StackOverflow, StarCoder) and mathematics (GSM8K), high-quality, fluent English prose often gets diluted when mixed with syntax-heavy code or rigid math… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smoltalk-creative-writing-enPurified-openai-messages.2026-09-10-delib-synth-smoke
Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7)
field
value
experiment
Deliberative SFT: delib; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7)
date_generated
20260910_191448
constitution
constitutions/abridged/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-10-delib-synth-smoke.details_bestdive__SmolLM3-3B-SFT-Free-Course
Smol course SFT evaluation - Kay Zheng
Actual full GSM8K test evaluation of bestdive/SmolLM3-3B-SFT-Free-Course, adapter revision 0484e028b494d605a267050a949c9266edadd16b, merged with pinned SmolLM3-3B-Base before evaluation.
Full 1319 test examples, zero-shot, original extractive_match: 0.4086429112964367 (stderr 0.013540639733342422).
Free Google Colab T4, no paid HF Jobs; cost 0.
lighteval 0.11.0, vLLM 0.10.1.1, Transformers 4.57.1, Python 3.12.
Dataset-address correction… See the full description on the dataset page: https://huggingface.co/datasets/bestdive/details_bestdive__SmolLM3-3B-SFT-Free-Course.test_smollm
MMLU-Pro Multi-Domain Dataset: test_smollm
Usage
from datasets import load_dataset
# Load entire dataset
dataset = load_dataset("dongboklee/test_smollm")
# Load specific domain
law_dataset = load_dataset("dongboklee/test_smollm", split="law")
2026-09-11-delib-sonnet-synth-smoke
Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7)
field
value
experiment
Deliberative SFT: delib-sonnet; native Qwen reasoning, best-of-2 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 2 runs >= 7)
date_generated
20260911_192737
constitution
constitutions/abridged/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-11-delib-sonnet-synth-smoke.lmsys-chat-1m-smortmodelsonlyThis version of the dataset only has responses from GPT-4, Claude-1, Claude-2, Claude-instant-1, and GPT-3.5-turbo
perfectblend-smoltalk-chinese-large-blend-regen2026-09-17-da-lowstakes-constitution-synth-smoke
18-row constitution-only low-stakes smoke; FAILED scaling gate; diagnostic candidates only
field
value
experiment
18-row constitution-only low-stakes smoke; FAILED scaling gate; diagnostic candidates only
date_generated
20260917_145731
constitution
constitutions/claude_distilled_09_principles/constitution.md sha256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc
source_repo
https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-da-lowstakes-constitution-synth-smoke.2026-09-17-da-lowstakes-values-in-advice-synth-smoke
18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only
field
value
experiment
18-row constitution-only low-stakes smoke; FAIL scaling gate; diagnostic candidates only
date_generated
20260917_171453
constitution
constitutions/claude_distilled_09_principles/constitution.md sha256 8e273b472d945aa23efa6236886da5e1171bff2193ee31ff73489ca54c4f0edc
source_repo
https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-da-lowstakes-values-in-advice-synth-smoke.2026-09-09-delib-synth-smoke
Deliberative SFT: delib; native Qwen reasoning, no judge filtering
field
value
experiment
Deliberative SFT: delib; native Qwen reasoning, no judge filtering
date_generated
20260909_152646
constitution
constitutions/abridged_no_delib/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 5302fa9589d32c4296d9393b4838a00d91042b76
models
qwen/qwen3.6-27b through Alibaba/OpenRouter (API revision not exposed)… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-09-delib-synth-smoke.qfs-smollm2-135m-wikitext2-native-v1
HF workflow d3dc69602aeb981f06bd9f4c726937f9
A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-native-bf16.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-native-v1.2026-09-17-delib-noref-synth-smoke
Deliberative SFT: delib-noref; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7)
field
value
experiment
Deliberative SFT: delib-noref; native Qwen reasoning, best-of-4 filtered by a constitution-aware judge (anthropic/claude-sonnet-5, min of 1 runs >= 7)
date_generated
20260917_191247
constitution
constitutions/abridged/constitution.md
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-17-delib-noref-synth-smoke.Generated-Empathetic-Dialogues-v0.1-Smol
Generated Empathetic Conversations v0.1 - Smol
This is dataset contains 10K rows of multi-round empathetic conversations convering a diverse set of topics.
Highlights
Multi-round conversation
It's not single-turn. The user and the assistant works together to gradually unfold the conversation.
The average number of turns is 5, with a standard deviation of approximately 1.59 turns. A turn consists of two messages with one by the user, and another by the… See the full description on the dataset page: https://huggingface.co/datasets/ychen/Generated-Empathetic-Dialogues-v0.1-Smol.smoke-openai-terra-batch-brasil-25-20260724-01
Smoke OpenAI Terra Batch — Brasil × 25 tasks
Run real de validação do fluxo matricial document_task_matrix, executada
sobre um único documento da Wikipédia em português com o título Brasil.
Cada uma das 25 tasks canônicas recebeu exatamente um slot inicial.
Resultado
status: completed
documentos: 1
pares planejados: 25
exemplos aceitos: 25
pares pulados: 0
pares esgotados: 0
resultados reais do backend: 27
retries com nova chamada: 2
backend: openai_api… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/smoke-openai-terra-batch-brasil-25-20260724-01.HuggingFaceTB__SmolLM-1.7B-Instruct-details
Dataset Card for Evaluation run of HuggingFaceTB/SmolLM-1.7B-Instruct
Dataset automatically created during the evaluation run of model HuggingFaceTB/SmolLM-1.7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceTB__SmolLM-1.7B-Instruct-details.2026-08-04-ddp-smoke-bundleDDP smoke bundle: train_lora.py + a 64-example toy set, for validating multi-GPU wiring.
2026-09-08-delib-synth-smoke
Deliberative SFT: delib; native Qwen reasoning, no judge filtering
field
value
experiment
Deliberative SFT: delib; native Qwen reasoning, no judge filtering
date_generated
20260908_182742
constitution
constitutions/no_claude_mentioned/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ a25750eb516cb6254640ae39a3a870f49398dbf0
models
qwen/qwen3.6-27b through Alibaba/OpenRouter (API revision not exposed)… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-08-delib-synth-smoke.smoltalk-binidxSmolThinker-Synthetic-Reasoning
SmolThinker Synthetic Reasoning Dataset
7,796 examples that teach a small instruct model to emit its reasoning inside
literal <think> ... </think> blocks, so chat UIs that render collapsible
reasoning (Open WebUI, Ollama, LM Studio) pick them up.
Model-agnostic: nothing in the data names a specific model. Reasoning length
scales with task difficulty, from one line on a greeting to 2,000+ characters
on a hard question.
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/SmolThinker-Synthetic-Reasoning.
