datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-stack-smol
Dataset Description
A small subset (~0.1%) of the-stack dataset, each programming language has 10,000 random samples from the original dataset. The dataset has 2.6GB of text (code).
Languages
The dataset contains 30 programming languages:
"assembly", "batchfile", "c++", "c", "c-sharp", "cmake", "css", "dockerfile", "fortran", "go", "haskell", "html", "java",
"javascript", "julia", "lua", "makefile", "markdown", "perl", "php", "powershell", "python", "ruby", "rust"… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol.smollm-corpus-cleaned
SmolLM-Corpus: Now shuffled and sharded (and Cleaned)!
This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming!
The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo.
Dataset Structure
The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.the-stack-smol-xl
Dataset Description
A small subset of the-stack dataset, with 87 programming languages, each has 10,000 random samples from the original dataset.
Languages
The dataset contains 87 programming languages:
'ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c',
'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir',
'elm', 'emacs-lisp','erlang'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol-xl.smol-worldcup
🏟️ Smol AI WorldCup — SHIFT Benchmark
The world's first 5-axis evaluation framework for small language models.
Not just "how smart?" — but "how honest? how fast? how small? how efficient?"
🏟️ Leaderboard
huggingface.co/spaces/ginigen-ai/smol-worldcup
📊 Dataset
huggingface.co/datasets/ginigen-ai/smol-worldcup
🏅 ALL Bench
huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard
🏆 Official Ranking: WCS (WorldCup Score)
WCS = √( SHIFT × PIR_norm )… See the full description on the dataset page: https://huggingface.co/datasets/ginigen-ai/smol-worldcup.smollm-corpus-fineweb-edu-enPurified-openai-messages
📖 smollm-corpus-fineweb-edu-enPurified-openai-messages
smollm-corpus-fineweb-edu-enPurified is a highly curated, "prose-first" subset of the fineweb-edu-dedup subset found in HuggingFaceTB/smollm-corpus.
The enPurified collection is built on a specific philosophy: Specialization. While the original dataset is excellent for general pre-training, high-quality fluent English prose often gets diluted when mixed with syntax-heavy code, rigid math formulas, or low-information web junk.… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-fineweb-edu-enPurified-openai-messages.smollm-corpus-cosmopedia-v2-enPurified-openai-messages
enPurified Collection: Smollm Corpus Cosmopedia V2]
Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows.
Purpose of the enPurified Collection
The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.smoltalk-chinese-QwQ-Distrill
smoltalk-chinese-QwQ-Distrill [中文] [English]
📖Technical Report
smoltalk-chinese-QwQ-Distrill is a Chinese fine-tuning dataset constructed with reference to the SmolTalk-Chinese dataset. It aims to provide high-quality synthetic reasoning data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese LLMs across various tasks… See the full description on the dataset page: https://huggingface.co/datasets/ChinaunicomSoftware/smoltalk-chinese-QwQ-Distrill.smoltalk-creative-writing-enPurified-openai-messages
📖 SmolTalk-Creative-Writing-enPurified-openai-messages
SmolTalk-Creative-Writing-enPurified is a highly curated, "prose-first" subset of the original collinear-ai/smoltalk-creative-writing dataset.
The enPurified collection is built on a specific philosophy: Specialization. While the ecosystem has plenty of datasets for coding (StackOverflow, StarCoder) and mathematics (GSM8K), high-quality, fluent English prose often gets diluted when mixed with syntax-heavy code or rigid math… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smoltalk-creative-writing-enPurified-openai-messages.Generated-Empathetic-Dialogues-v0.1-Smol
Generated Empathetic Conversations v0.1 - Smol
This is dataset contains 10K rows of multi-round empathetic conversations convering a diverse set of topics.
Highlights
Multi-round conversation
It's not single-turn. The user and the assistant works together to gradually unfold the conversation.
The average number of turns is 5, with a standard deviation of approximately 1.59 turns. A turn consists of two messages with one by the user, and another by the… See the full description on the dataset page: https://huggingface.co/datasets/ychen/Generated-Empathetic-Dialogues-v0.1-Smol.smoke-openai-terra-batch-brasil-25-20260724-01
Smoke OpenAI Terra Batch — Brasil × 25 tasks
Run real de validação do fluxo matricial document_task_matrix, executada
sobre um único documento da Wikipédia em português com o título Brasil.
Cada uma das 25 tasks canônicas recebeu exatamente um slot inicial.
Resultado
status: completed
documentos: 1
pares planejados: 25
exemplos aceitos: 25
pares pulados: 0
pares esgotados: 0
resultados reais do backend: 27
retries com nova chamada: 2
backend: openai_api… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/smoke-openai-terra-batch-brasil-25-20260724-01.SmolThinker-Synthetic-Reasoning
SmolThinker Synthetic Reasoning Dataset
7,796 examples that teach a small instruct model to emit its reasoning inside
literal <think> ... </think> blocks, so chat UIs that render collapsible
reasoning (Open WebUI, Ollama, LM Studio) pick them up.
Model-agnostic: nothing in the data names a specific model. Reasoning length
scales with task difficulty, from one line on a greeting to 2,000+ characters
on a hard question.
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/SmolThinker-Synthetic-Reasoning.smoltalk2_everyday_conv_pt
SMOL Everyday Conversation PT
This dataset consists of a Portuguese version of the smoltalk_smollm3_everyday_conversations_no_think split of HuggingFaceTB/smoltalk2. The first two user turns were translated as well as the first assistant turn, the continuation of the conversation was generated using Gemma 3-27B.
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smoltalk2_everyday_conv_pt.smoothie-qwen3-8b-kr-self-driving-legal-dataset-v5-cot
🇰🇷 자율주행법령 CoT 파인튜닝 데이터셋 v5
왜 이 데이터셋을 새로 만들었는가?
기존 dataset-v3 는 단순 질문-답변(Direct To Response, DTRO Style) 포맷으로 구성되어 있었습니다.
// v3 포맷 (기존)
{
"instruction": "자율주행자동차란 무엇인가요?",
"output": "자율주행자동차란 ..."
}
이 방식으로 파인튜닝한 모델(v3)을 RAG 파이프라인과 결합하여 평가한 결과, 정답률 43% 로 순정 모델(90%)에 크게 뒤처지는 것이 확인되었습니다. 실패의 핵심 원인은 다음과 같습니다:
실패 원인
설명
템플릿 과적합
모델이 논리가 아닌 답변 패턴(아닙니다 + 설명)을 암기
RAG 컨텍스트 무시
학습된 내부 패턴이 외부 검색 문서를 압도
<think> 태그 미사용
Qwen3의 추론(Chain-of-Thought) 능력이 전혀 활성화되지 않음… See the full description on the dataset page: https://huggingface.co/datasets/bluejude10/smoothie-qwen3-8b-kr-self-driving-legal-dataset-v5-cot.empathetic_dialogues_ko
Dataset Card for "한국어 일상 속 공감형 대화 데이터셋(멀티-턴)"
Dataset Summary
boostCamp AI Tech 5기 과정 중 NLP 12조 훈제연어들 팀의 최종 프로젝트에서 제작한 데이터입니다.
일상 속 다양한 상황에서 사용자와 챗봇 간의 대화를 담은 데이터셋 입니다.
GPT4, GPT3.5-turbo로 제작된 합성데이터이며 싱글-턴, 2-턴, 3-턴 대화로 구성되어 있습니다.
답변은 [공감적 표현 - 일반적인 대화 - 관련된 질문] 의 형태를 가집니다.
Generation Prompt Example(GPT3.5-turbo)
Take a close look at the following example and Conditions. Create nine sessions that each of the session is ongoing conversation about a single… See the full description on the dataset page: https://huggingface.co/datasets/Smoked-Salmon-s/empathetic_dialogues_ko.prism-smopPRISM-SMOP
Датасет генераций кода 1С:Предприятие (BSL) с четырёхосевой разметкой качества по метрике SMOP
🤗 Hugging Face ·
Бенчмарк PRISM ·
Лицензия CC BY 4.0 · Метрика SMOP
1225 генераций кода 1С:Предприятие (BSL) от 35 нейросетей на 35 задачах бенчмарка PRISM, размеченных по четырём осям метрики SMOP — синтаксис, смысл, оптимальность, платформа — автоматическим оценщиком L1. Внутри — SFT-подвыборка из 222 решений открытых моделей, прошедших все скрытые тесты.
1225 BSL… See the full description on the dataset page: https://huggingface.co/datasets/genlab-1c/prism-smop.smol-rewrite-PT
SMOL Rewrite PT
This dataset is the translated version of the smol-rewrite subset of the HuggingFaceTB/smoltalk.
This dataset includes an high-quality split used in the ramp down phase of the AMALIA's model post-training. The quality classification was done using google/gemma-4-31B-it.
Original Dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk
Note: This dataset comprises machine translated content and may contain translation errors or artifacts.
This… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smol-rewrite-PT.smol_summarize_pt
SMOL Summarize PT
This dataset is the translated version of the smol-summarize subset of the HuggingFaceTB/smoltalk.
Original Dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk
Note: This dataset comprises machine translated content and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model.
Citation
If you use this dataset or… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smol_summarize_pt.smollm3-3b-base-blind-spots
SmolLM3-3B-Base Blind Spots Dataset
This dataset contains 10 test cases where I explored the failure modes of
SmolLM3-3B-Base,
a 3 billion parameter base language model released by HuggingFace in 2025.
The goal was to find diverse cases where the model makes clearly incorrect
or unexpected completions its "blind spots."
Model Tested
Model: HuggingFaceTB/SmolLM3-3B-Base
Parameters: 3B
Type: Base pretrained model
License: Apache 2.0
How I Loaded the Model
I… See the full description on the dataset page: https://huggingface.co/datasets/FatimaAfzal01/smollm3-3b-base-blind-spots.numina_smoltalk_mixturegemma-3n-4b-distill-smollm2-360m-instruct-425xTrace of Gemma 3n 4B Distill SmolLM2 360M Instruct LLM by sapbot (me).
Data count (Total: 425):
English - 209
Russian - 216
Data is presented in ShareGPT format and each conversation split by newline.
Note: This was added more as a "examples" of this model's outputs. Of course you will not distill a distilled model (I hope).
Brought to you by sapbot from Romarchive
matilda-smollm-mix-15b-gpt2
matilda-smollm-mix-15B-gpt2
15 B GPT-2-BPE tokens drawn from a 5:1 token-balanced mix of
HuggingFaceTB/smollm-corpus:
Source
Share
Tokens
fineweb-edu-dedup
83.33 %
12.50 B
cosmopedia-v2
16.67 %
2.50 B
Total: 15,000,349,569 tokens across 151 shards (shard_*.bin, uint16,
100 M tokens per shard).
The full SmolLM recipe is 75 / 15 / 10 fineweb-edu / cosmopedia-v2 / python-edu.
python-edu was dropped because the HuggingFaceTB/smollm-corpus subset
ships only blob_id… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/matilda-smollm-mix-15b-gpt2.smooth_reading
Smooth Reading
This dataset accompanies the paper Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Understanding.
license: mit
team-coop-smoke
CooperBench Team → Coop (Qwen3.5-9B smoke)
Placeholder / smoke dataset (1 trajectory). A single 2-agent
cooperbench team run (lead + member, no protocol) reshaped into the
2-agent coop layout defined in cooperbench/CooperData PR
#98.
The full dataset is the canonical place where future team→coop conversions
will land; this entry validates the converter and the publishing pipeline.
Source
Source run
logs/qwen35-smoke-mini-team-noproto/
Source repo… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/team-coop-smoke.balanced_smoltalk_basesyngen-reasoning-example-80-smoltalk1Reasoning generated using https://huggingface.co/Pinkstack/syngen-reasoning-0.6b (some of it was cutt off due to 8192 max length instead of 16k)
based on the original smoltalk
smoll-cotThe dataset its still under development
smollm3-3b-base-blind-spots
SmolLM3-3B-Base Blind Spots
Title & Overview
A curated set of failure cases for HuggingFaceTB/SmolLM3-3B-Base, showcasing blind spots discovered while probing the 3B-parameter base pre-training checkpoint released in July 2025. Each entry captures a prompt, the expected aligned behaviour, and the model's actual output. The dataset illustrates common failure patterns observed when probing the base model without any instruction tuning, RLHF, or safety fine-tuning applied.… See the full description on the dataset page: https://huggingface.co/datasets/aneeshadas02/smollm3-3b-base-blind-spots.smollm3-blindspots
Blind Spots of SmolLM3-3B-Base
This dataset documents systematic failure cases ("blind spots") observed
when evaluating the SmolLM3-3B-Base model.
The goal of this dataset is to identify patterns where a small base
language model struggles with reasoning tasks that require precise
symbolic or character-level manipulation.
The dataset contains prompts where the model produces incorrect answers
compared to the expected output.
Model Tested
Model:… See the full description on the dataset page: https://huggingface.co/datasets/hans1337/smollm3-blindspots.smollm2-dpo-preferences
DPO Preferences Dataset (Restricted Access)
Access Policy (Restricted)
This dataset repo is public with manual gated access.
Only approved users (from lums.edu.pk) will be granted access.
Intended Use
Preference optimization / DPO experiments for model alignment.
Research and controlled evaluation.
Out-of-Scope Use
Any harmful, abusive, or policy-violating application.
Safety-critical deployment without additional safeguards.
Files… See the full description on the dataset page: https://huggingface.co/datasets/Amin-AQ/smollm2-dpo-preferences.
