datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smollm-corpus
SmolLM-Corpus
This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models.
You can find more details about the models trained on this dataset in our SmolLM blog post.
Dataset subsets
Cosmopedia v2
Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus.smollm-chunked
FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora
This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity.
Full Documentation
For complete usage instructions, installation guide, and tutorial, please refer to:
Main Tutorial README
Data Distribution
Due to Hugging Face storage quota… See the full description on the dataset page: https://huggingface.co/datasets/enguyen/smollm-chunked.smollm-12.5-corpus
SmolLM-1/8-Corpus
Around 1/8 upper-quality subset of SmolLM Corpus for training Chinchilla-optimal GPT-2 scale (sub 1.5B) models, which is a good scale for verifying a model architecture under the scaling laws.
Firstly filtered samples with int_score >=4 from FineWeb-edu-dedup, then keep the training mixture with the same distribution from SmolLM.
In which FineWeb-Edu-dedup occupies around 70% of the corpus. Then sample other dataset based on the mixture ratios respectively. For… See the full description on the dataset page: https://huggingface.co/datasets/chengjunyan1/smollm-12.5-corpus.seq2seq-mixed-pretraining-SmolLM2jupyter-scripts-smollm3
The Stack v2 Jupyter Notebooks as Scripts
This dataset contains script representations of the Jupyter notebooks in
The Stack v2. It was
created from the materialized Jupyter_Notebook split in
jordangong/the-stack-v2-smollm3.
The output schema follows the Jupyter-script schema used by
bigcode/starcoderdata,
but this release is not deduplicated, PII-filtered, or otherwise equivalent
to StarCoderData's filtered split.
Relationship to the SmolLM3 training mix
This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.smollm-10smollm-corpus-3.5M
A very small version of smollm-corpus (+finemath-4plus) for experimenting llm pre-training.
cosmopedia-v2 (1M rows)
fineweb-edu-dedup (1M rows)
python-edu (0.5M rows)
finemath-4plus (1M rows)
smollm-corpus
SmolLM-Corpus
This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models.
You can find more details about the models trained on this dataset in our SmolLM blog post.
Dataset subsets
Cosmopedia v2
Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/oieieio/smollm-corpus.smollm2-135m-instruct-SAE
Layer
EV
Mean L0
Recon Loss
Dead %
0
0.9480
48.74
0.2074
0.0
1
0.9599
43.65
0.3298
0.0
2
0.9631
46.81
0.5021
0.0
3
0.9508
46.56
0.7462
0.0
4
0.9463
46.23
0.8936
0.0
5
0.9350
47.57
1.1605
0.0
6
0.9306
48.44
1.3838
0.0
7
0.9318
49.51
1.5446
0.0
8
0.9432
46.52
1.6598
0.0
9
0.9373
47.15
2.0706
0.0
10
0.9348
45.53
2.2983
0.0
11
0.9905
48.58
5.8113
0.0
12
0.9901
48.42
6.1039
0.0
13
0.9891
46.15
6.9692
0.0
14
0.9884
44.76
7.1844
0.0
15
0.9863
47.63
8.6521
0.0… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/smollm2-135m-instruct-SAE.Bitext-SmolLM2-1024-natural-instructions-formatai_human_smollm360m_logits
AI/Human logits — SmolLM-360M
Next-token distributions from HuggingFaceTB/SmolLM-360M over
AI_Human_Dataset.csv from Aanimated/telescope_datasets.
At each position, only tokens with softmax probability >= the cutoff are
kept (the argmax is always retained), sorted by descending probability.
Columns
column
description
sample_index
index of the source row
input_ids
SmolLM token ids for the sample
num_tokens
sequence length after truncation
token_ids… See the full description on the dataset page: https://huggingface.co/datasets/Joshfcooper/ai_human_smollm360m_logits.smollm3-tracesportuguese-eval-logs-olmo2-smollm3
Evaluation Logs on Portuguese Benchmarks for OLMo-2 and SmolLM3
These logs contain benchmark results across a suite of Portuguese-language tasks. The data consists of recordings of the performance of various 3 different models at different checkpoints throughout their pretraining runs:
SmolLM3
OLMo-2-0425-1B
OLMo-2-1124-7B
Splits
Each split (smollm3_3b, olmo2_1b, olmo2_7b) contains rows for model checkpoints and columns for benchmark scores (e.g., ASSIN2 RTE, ENEM… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-eval-logs-olmo2-smollm3.pretraining-pretokenized-smollm3
SmolLM3 Pretokenized Pretraining Sources
Datatrove/Nanotron tokenized-byte versions of three public pretraining sources:
fineweb-edu-10bt: HuggingFaceFW/fineweb-edu, sample/10BT
finemath-4plus: HuggingFaceTB/finemath, finemath-4plus
stack-edu-python: HuggingFaceTB/stack-edu, Python, with content retrieved from the public Software Heritage S3 bucket using the upstream dataset-card procedure
All subsets use HuggingFaceTB/SmolLM3-3B at revision… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/pretraining-pretokenized-smollm3.RULER-262144-SmolLM3-11T-32k-v1-remote-codeRULER-131072-SmolLM3-11T-32k-v1-remote-codevery-smollm-corpus-0.5Mqfs-smollm2-135m-wikitext2-native-v1
HF workflow d3dc69602aeb981f06bd9f4c726937f9
A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-native-bf16.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-native-v1.HuggingFaceTB__SmolLM-1.7B-Instruct-details
Dataset Card for Evaluation run of HuggingFaceTB/SmolLM-1.7B-Instruct
Dataset automatically created during the evaluation run of model HuggingFaceTB/SmolLM-1.7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceTB__SmolLM-1.7B-Instruct-details.smollm-corpus
SmolLM-Corpus
This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models.
You can find more details about the models trained on this dataset in our SmolLM blog post.
Dataset subsets
Cosmopedia v2
Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/hhoenjet/smollm-corpus.qfs-smollm2-135m-wikitext2-gptq-g32-v1
HF workflow 32c6ab05b0ceab1cecdceda838846388
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-gptq-int4-g32.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-gptq-g32-v1.RULER-65536-SmolLM3-11T-32k-v1-remote-codeqfs-smollm2-135m-wikitext2-gptq-g64-v1
HF workflow 73f0a12a901c7368794a3a886f55b675
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-gptq-int4-g64.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-gptq-g64-v1.RULER-4096-SmolLM3-11T-32k-v1-remote-codesmollm-corpus
SmolLM-Corpus
This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models.
You can find more details about the models trained on this dataset in our SmolLM blog post.
Dataset subsets
Cosmopedia v2
Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/Mgmgrand420/smollm-corpus.meditsolutions__SmolLM2-MedIT-Upscale-2B-details
Dataset Card for Evaluation run of meditsolutions/SmolLM2-MedIT-Upscale-2B
Dataset automatically created during the evaluation run of model meditsolutions/SmolLM2-MedIT-Upscale-2B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meditsolutions__SmolLM2-MedIT-Upscale-2B-details.FlofloB__smollm2-135M_pretrained_1000k_fineweb-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1000k_fineweb
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1000k_fineweb
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1000k_fineweb-details.FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_selected-details
Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_selected
Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_selected
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_selected-details.details_pmakiela__SmolLM3-3B-dpo-v0_1_private
Dataset Card for Evaluation run of pmakiela/SmolLM3-3B-dpo-v0_1
Dataset automatically created during the evaluation run of model pmakiela/SmolLM3-3B-dpo-v0_1.
The dataset is composed of 4 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/pmakiela/details_pmakiela__SmolLM3-3B-dpo-v0_1_private.compilation-SmolLM2
