datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-stack-v2-smollm3
The Stack v2 — materialized source code
Upstream dataset:
bigcode/the-stack-v2
Exact upstream commit:
e565caa3a78c2423bd374333a472b049eb090e47
Primary source-content endpoint:
https://softwareheritage.s3.amazonaws.com/content/{blob_id}
Configurations
TypeScript
Swift
Ruby
Rust
Go
Shell
Jupyter_Notebook
HTML
Python
Java
JavaScript
C
C++
C-Sharp
PHP
SQL
Markdown
Added columns
content: decoded source content
download_error: null on successful… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/the-stack-v2-smollm3.smollm-corpus-cleaned
SmolLM-Corpus: Now shuffled and sharded (and Cleaned)!
This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming!
The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo.
Dataset Structure
The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.jupyter-scripts-smollm3
The Stack v2 Jupyter Notebooks as Scripts
This dataset contains script representations of the Jupyter notebooks in
The Stack v2. It was
created from the materialized Jupyter_Notebook split in
jordangong/the-stack-v2-smollm3.
The output schema follows the Jupyter-script schema used by
bigcode/starcoderdata,
but this release is not deduplicated, PII-filtered, or otherwise equivalent
to StarCoderData's filtered split.
Relationship to the SmolLM3 training mix
This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.apigen-smollm-trl-FC
Dataset card for argilla-warehouse/apigen-smollm-trl-FC
This dataset is a merge of argilla/Synth-APIGen-v0.1
and Salesforce/xlam-function-calling-60k, and was prepared for training using the script
prepare_for_sft.py that can be found in the repository files.
References
@article{liu2024apigen,
title={APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets},
author={Liu, Zuxin and Hoang, Thai and Zhang, Jianguo and Zhu, Ming and… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/apigen-smollm-trl-FC.smollm-corpus-fineweb-edu-enPurified-openai-messages
📖 smollm-corpus-fineweb-edu-enPurified-openai-messages
smollm-corpus-fineweb-edu-enPurified is a highly curated, "prose-first" subset of the fineweb-edu-dedup subset found in HuggingFaceTB/smollm-corpus.
The enPurified collection is built on a specific philosophy: Specialization. While the original dataset is excellent for general pre-training, high-quality fluent English prose often gets diluted when mixed with syntax-heavy code, rigid math formulas, or low-information web junk.… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-fineweb-edu-enPurified-openai-messages.smollm-corpus-cosmopedia-v2-enPurified-openai-messages
enPurified Collection: Smollm Corpus Cosmopedia V2]
Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows.
Purpose of the enPurified Collection
The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming
A corpus of high quality fine tuning data meant for fine tuning various HelixLM models
Dataset Composition:
A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ...
Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning.
Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.ai-vs-human-HuggingFaceTB-SmolLM2-1.7B-Instruct
AI vs Human dataset on the CNN Daily mails
Dataset Description
This dataset showcases pairs of truncated articles and their respective completions, crafted either by humans or an AI language model.
Each article was randomly truncated between 25% and 50% of its length.
The language model was then tasked with generating a completion that mirrored the characters count of the original human-written continuation.
Data Fields
'human': The original human-authored… See the full description on the dataset page: https://huggingface.co/datasets/zcamz/ai-vs-human-HuggingFaceTB-SmolLM2-1.7B-Instruct.1b-smollm-corpus
SmolLM-Corpus — 1B Token Subset
A curated 1-billion-token English pretraining corpus sampled from
HuggingFaceTB/smollm-corpus,
designed for training small language models (~20M parameters).
Dataset Composition
Source
Ratio
Tokens
Documents
FineWeb-Edu (dedup)
87%
~870M
849,577
Cosmopedia v2
13%
~130M
161,889
Total
100%
~1B
1,011,466
Rationale for the Split
The 87/13 ratio mirrors the natural token distribution of the full… See the full description on the dataset page: https://huggingface.co/datasets/ecreeth/1b-smollm-corpus.traffic-accidents-reports-kd-smollm2-360M-7k
Accident Reporting KD Dataset (One-Paragraph)
Short description.A training/evaluation dataset for generating one-paragraph accident/incident reports from structured facts.This dataset mixes gold human targets from zBotta/traffic-accidents-reports-5k with teacher-generated soft targets produced by the model zBotta/smollm2-accident-reporter-360m to support knowledge distillation (KD) of a smaller student.
Output style: a single paragraph, neutral tone, covering What, When, Where, Who… See the full description on the dataset page: https://huggingface.co/datasets/DSTI/traffic-accidents-reports-kd-smollm2-360M-7k.SmoLLM-Dataset
Dataset Card for SmoLLM-Dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/KBaba7/SmoLLM-Dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/KBaba7/SmoLLM-Dataset.smollm3-3b-base-blind-spots
SmolLM3-3B-Base Blind Spots Dataset
This dataset contains 10 test cases where I explored the failure modes of
SmolLM3-3B-Base,
a 3 billion parameter base language model released by HuggingFace in 2025.
The goal was to find diverse cases where the model makes clearly incorrect
or unexpected completions its "blind spots."
Model Tested
Model: HuggingFaceTB/SmolLM3-3B-Base
Parameters: 3B
Type: Base pretrained model
License: Apache 2.0
How I Loaded the Model
I… See the full description on the dataset page: https://huggingface.co/datasets/FatimaAfzal01/smollm3-3b-base-blind-spots.smollm2-135m-abstention-posttrain-data
SmolLM2-135M abstention post-training task
Synthetic abstention-vs-fabrication task used in
barlowa124/llm-posttraining
and the checkpoints at
barlowa/smollm2-135m-abstention-posttrain.
Files
data/ — the four training/eval splits. sft (960 prompt+completion),
dpo (400 chosen/rejected pairs), eval (160 held-out entities),
rl (400 GRPO prompts). Held-out entities never appear in train —
contamination-free by construction.
responses/ — raw model generations per… See the full description on the dataset page: https://huggingface.co/datasets/barlowa/smollm2-135m-abstention-posttrain-data.ai-vs-human-HuggingFaceTB-SmolLM2-360M-Instruct
AI vs Human dataset on the CNN Daily mails
Dataset Description
This dataset showcases pairs of truncated articles and their respective completions, crafted either by humans or an AI language model.
Each article was randomly truncated between 25% and 50% of its length.
The language model was then tasked with generating a completion that mirrored the characters count of the original human-written continuation.
Data Fields
'human': The original human-authored… See the full description on the dataset page: https://huggingface.co/datasets/zcamz/ai-vs-human-HuggingFaceTB-SmolLM2-360M-Instruct.matilda-smollm-mix-15b-gpt2
matilda-smollm-mix-15B-gpt2
15 B GPT-2-BPE tokens drawn from a 5:1 token-balanced mix of
HuggingFaceTB/smollm-corpus:
Source
Share
Tokens
fineweb-edu-dedup
83.33 %
12.50 B
cosmopedia-v2
16.67 %
2.50 B
Total: 15,000,349,569 tokens across 151 shards (shard_*.bin, uint16,
100 M tokens per shard).
The full SmolLM recipe is 75 / 15 / 10 fineweb-edu / cosmopedia-v2 / python-edu.
python-edu was dropped because the HuggingFaceTB/smollm-corpus subset
ships only blob_id… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/matilda-smollm-mix-15b-gpt2.smollm_self_data
smollm_self_data
Dataset Description
This dataset consists of 100 question-answer pairs generated entirely by the HuggingFaceTB/SmolLM2-135M-Instruct model. The dataset was created using a "self-prompting" approach, where the model was first asked to generate an interesting question, and then asked to provide an answer to that same question.
The primary goal of this dataset is to serve as a base for fine-tuning of other small conversational models.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Fu01978/smollm_self_data.smollm2-blind-spots
SmolLM2-1.7B Blind Spots Dataset
This dataset contains 10 diverse examples where the SmolLM2-1.7B base model makes incorrect predictions or demonstrates "blind spots".
Model Tested
Model: SmolLM2-1.7B
Parameters: 1.7 Billion
Type: Base (Pre-trained)
How to Load the Model
The model was loaded using the transformers library in Python.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "HuggingFaceTB/SmolLM2-1.7B"
tokenizer =… See the full description on the dataset page: https://huggingface.co/datasets/Alhibb/smollm2-blind-spots.gemma-3n-4b-distill-smollm2-360m-instruct-425xTrace of Gemma 3n 4B Distill SmolLM2 360M Instruct LLM by sapbot (me).
Data count (Total: 425):
English - 209
Russian - 216
Data is presented in ShareGPT format and each conversation split by newline.
Note: This was added more as a "examples" of this model's outputs. Of course you will not distill a distilled model (I hope).
Brought to you by sapbot from Romarchive
smollm3-base-blind-spots
SmolLM3-3B-Base — Blind Spots Dataset
Overview
This dataset documents 12 diverse blind spots of the base language model
HuggingFaceTB/SmolLM3-3B-Base
(3 billion parameters, Apache-2.0, released 2025).
Each row contains:
Field
Description
id
Integer index
category
Type of reasoning required
input_prompt
Partial text fed to the model
expected_output
Correct continuation
model_output
What SmolLM3-3B-Base actually generated
notes
Explanation of why the… See the full description on the dataset page: https://huggingface.co/datasets/NiazTahi/smollm3-base-blind-spots.smollm3-3b-base-blind-spots
SmolLM3-3B-Base Blind Spots
Title & Overview
A curated set of failure cases for HuggingFaceTB/SmolLM3-3B-Base, showcasing blind spots discovered while probing the 3B-parameter base pre-training checkpoint released in July 2025. Each entry captures a prompt, the expected aligned behaviour, and the model's actual output. The dataset illustrates common failure patterns observed when probing the base model without any instruction tuning, RLHF, or safety fine-tuning applied.… See the full description on the dataset page: https://huggingface.co/datasets/aneeshadas02/smollm3-3b-base-blind-spots.smollm3-3b-base-blindspots
SmolLM3-3B-Base — Blind Spots Dataset
This dataset contains 10 diverse input-output pairs where the base language model
HuggingFaceTB/SmolLM3-3B-Base
produces incorrect predictions under greedy decoding. Each row records the exact prompt fed
to the model, the correct expected answer, and what the model actually generated — along with
a description of the error type.
Model Tested
Field
Value
Model
HuggingFaceTB/SmolLM3-3B-Base
Parameters
3 billion… See the full description on the dataset page: https://huggingface.co/datasets/Dhruba461/smollm3-3b-base-blindspots.smollm2-1.7b-blind-spots
SmolLM2-1.7B Blind Spots Dataset
A curated evaluation dataset documenting specific failure modes of HuggingFaceTB/SmolLM2-1.7B — a 1.7 billion parameter pretrained (base) language model. Each entry contains a completion-style prompt, the verified correct answer, and the model's actual incorrect output produced via deterministic greedy decoding.
This dataset was created as part of the "Blind Spots of Frontier Models" technical challenge to systematically identify where small… See the full description on the dataset page: https://huggingface.co/datasets/Shinzmann/smollm2-1.7b-blind-spots.smoll-cotThe dataset its still under development
smollm2-1.7b-blindspots
SmolLM2-1.7B Blind Spots Dataset
A curated dataset of 12 diverse probe examples where the base language model
HuggingFaceTB/SmolLM2-1.7B
makes incorrect or unreliable predictions. Each row contains the raw prompt, the
expected correct answer, the model's actual output (greedy decoding), the error
category, and an explanation.
Tested Model
HuggingFaceTB/SmolLM2-1.7B
Property
Value
Parameters
1.7 billion
Type
Pure base model (pretrained only — no… See the full description on the dataset page: https://huggingface.co/datasets/Shah-4-8-1-2/smollm2-1.7b-blindspots.smollm2-blind-spots
Model Tested
HuggingFaceTB/SmolLM2-1.7B
How I loaded it
Used HuggingFace Transformers with AutoModelForCausalLM on Google Colab (T4 GPU, float16). Greedy decoding (do_sample=False) for reproducibility.
View Colab Notebook
Blind Spots Found
The model struggled with: multi-step arithmetic, low-resource languages (Yoruba), African geographic knowledge, code generation, and logical reasoning.
Fine-tuning Dataset Recommendation
GSM8K / MATH for… See the full description on the dataset page: https://huggingface.co/datasets/mirackchuks/smollm2-blind-spots.smollm3-base-blindspots
SmolLM3-3B-Base Blind Spots Evaluation Dataset
Dataset Summary
This dataset documents 10 diverse failure cases discovered while evaluating
HuggingFaceTB/SmolLM3-3B-Base,
a 3-billion parameter decoder-only base language model released by Hugging Face in July 2025.
The evaluation was conducted as part of the Fatima Fellowship technical challenge on Blind Spots of Frontier Models.
Model Tested
Model: HuggingFaceTB/SmolLM3-3B-Base
Parameters: 3 billion… See the full description on the dataset page: https://huggingface.co/datasets/habibahabchi/smollm3-base-blindspots.smollm2-blindspots
Blind Spots Dataset for SmolLM2-1.7B
Model Tested
Model: SmolLM2-1.7B
Parameters: 1.7B
Release Date: February 2025
How I Loaded the Model
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "HuggingFaceTB/SmolLM2-1.7B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
device_map="auto",
torch_dtype="auto"
)
def test_model(prompt, max_new_tokens=100):
inputs =… See the full description on the dataset page: https://huggingface.co/datasets/Muhammad0981/smollm2-blindspots.smollm3-blindspots
Blind Spots of SmolLM3-3B-Base
This dataset documents systematic failure cases ("blind spots") observed
when evaluating the SmolLM3-3B-Base model.
The goal of this dataset is to identify patterns where a small base
language model struggles with reasoning tasks that require precise
symbolic or character-level manipulation.
The dataset contains prompts where the model produces incorrect answers
compared to the expected output.
Model Tested
Model:… See the full description on the dataset page: https://huggingface.co/datasets/hans1337/smollm3-blindspots.smollm3-blind-spots
SmolLM3-3B-Base Blind Spots
This dataset documents 10 failure cases observed while testing the model HuggingFaceTB/SmolLM3-3B-Base in Google Colab.
Model Tested
Model: HuggingFaceTB/SmolLM3-3B-Base
How the Model Was Loaded
The model was loaded in Google Colab using the transformers library with 4-bit quantization to run on limited GPU resources.
Dataset Description
This dataset contains 10 examples where the model produced incorrect outputs or failed… See the full description on the dataset page: https://huggingface.co/datasets/Sgobir/smollm3-blind-spots.smollm-corpus-python
smollm-corpus - python
A version of the python-edu subset with the text added
