datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-stack-v2-smollm3
The Stack v2 — materialized source code
Upstream dataset:
bigcode/the-stack-v2
Exact upstream commit:
e565caa3a78c2423bd374333a472b049eb090e47
Primary source-content endpoint:
https://softwareheritage.s3.amazonaws.com/content/{blob_id}
Configurations
TypeScript
Swift
Ruby
Rust
Go
Shell
Jupyter_Notebook
HTML
Python
Java
JavaScript
C
C++
C-Sharp
PHP
SQL
Markdown
Added columns
content: decoded source content
download_error: null on successful… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/the-stack-v2-smollm3.jupyter-scripts-smollm3
The Stack v2 Jupyter Notebooks as Scripts
This dataset contains script representations of the Jupyter notebooks in
The Stack v2. It was
created from the materialized Jupyter_Notebook split in
jordangong/the-stack-v2-smollm3.
The output schema follows the Jupyter-script schema used by
bigcode/starcoderdata,
but this release is not deduplicated, PII-filtered, or otherwise equivalent
to StarCoderData's filtered split.
Relationship to the SmolLM3 training mix
This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.smollm3-configs
SmolLM3 Training Configs
[IMPORTANT NOTE]: for the latest configs go to this repo: https://github.com/huggingface/smollm/tree/main/text/pretraining/smollm3
Here you can find the training configs for SmoLLM3-3B-Base using nanotron with exact training details and data mixtures.
The model was trained on 11.2T tokens in 3 stages on 4k context:
stage 1 config
stage 2 config
stage 3 config
And then we trained on an additional 2 stages to extend the contetx length to 64k:
stage 4… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/smollm3-configs.c4-rewritten-14b-retok-smollm360msmollm3-baseline_v3dclm-14b-c4-rewritten-14b-retok-smollm360mdclm-6.7b-c4-rewritten-6.7b-retok-smollm360msmollm3-blueprintHere you can find the SmolLM3 Engineering Blueprint
c4-rewritten-6.7b-retok-smollm360mpretraining-pretokenized-smollm3
SmolLM3 Pretokenized Pretraining Sources
Datatrove/Nanotron tokenized-byte versions of three public pretraining sources:
fineweb-edu-10bt: HuggingFaceFW/fineweb-edu, sample/10BT
finemath-4plus: HuggingFaceTB/finemath, finemath-4plus
stack-edu-python: HuggingFaceTB/stack-edu, Python, with content retrieved from the public Software Heritage S3 bucket using the upstream dataset-card procedure
All subsets use HuggingFaceTB/SmolLM3-3B at revision… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/pretraining-pretokenized-smollm3.smollm3-wikirollouts-smollm3-rl
rollouts-smollm3-rl — evaluation rollouts (boxed prompt)
Model: reasoning-cues/smollm3-rl @ 97bda61 — a GRPO LoRA adapter (r=64, alpha=128, all attention and
MLP projections) on HuggingFaceTB/SmolLM3-3B-Base @ d78a42f, merged with peft merge_and_unload
(bf16) and served with the base model's tokenizer files (the adapter repo's tokenizer_config.json
declares TokenizersBackend; its tokenizer.json is byte-identical to the base's). Protocol: the paper's
released-model grid — 32… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-smollm3-rl.ai_human_smollm360m_logits
AI/Human logits — SmolLM-360M
Next-token distributions from HuggingFaceTB/SmolLM-360M over
AI_Human_Dataset.csv from Aanimated/telescope_datasets.
At each position, only tokens with softmax probability >= the cutoff are
kept (the argmax is always retained), sorted by descending probability.
Columns
column
description
sample_index
index of the source row
input_ids
SmolLM token ids for the sample
num_tokens
sequence length after truncation
token_ids… See the full description on the dataset page: https://huggingface.co/datasets/Joshfcooper/ai_human_smollm360m_logits.smollm3-baseline_v2smollm3-tracessmollm3-infiwebmathsmollm3-github-issuessmollm3-stackexchangeportuguese-eval-logs-olmo2-smollm3
Evaluation Logs on Portuguese Benchmarks for OLMo-2 and SmolLM3
These logs contain benchmark results across a suite of Portuguese-language tasks. The data consists of recordings of the performance of various 3 different models at different checkpoints throughout their pretraining runs:
SmolLM3
OLMo-2-0425-1B
OLMo-2-1124-7B
Splits
Each split (smollm3_3b, olmo2_1b, olmo2_7b) contains rows for model checkpoints and columns for benchmark scores (e.g., ASSIN2 RTE, ENEM… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-eval-logs-olmo2-smollm3.smollm3-stack-v2-TypeScriptRULER-131072-SmolLM3-11T-32k-v1-remote-codesmollm3-stack-v2-Javarollouts-smollm3-csail
rollouts-smollm3-csail
Model: HuggingFaceTB/SmolLM3-3B-Base and its intermediate checkpoints (smollm3_ladder/).
Tokenizer: HuggingFaceTB/SmolLM3-3B-Base.
Protocol: entropy screens: none + the model's top-20 beam nominees, first 30 MATH-train problems x 16 rollouts, budget 16,384, T 0.6, top-p 0.95, seed 20260819; code screens: MBPP beam nominees on HumanEval 164 x 32, budget 31,744, execution-graded copies included; smollm3_confirm: MATH-500 x 4 at 31,744 (none, p2_bare, p2_okay… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-smollm3-csail.RULER-131072-SmolLM3-SFTdetails_bestdive__SmolLM3-3B-SFT-Free-Course
Smol course SFT evaluation - Kay Zheng
Actual full GSM8K test evaluation of bestdive/SmolLM3-3B-SFT-Free-Course, adapter revision 0484e028b494d605a267050a949c9266edadd16b, merged with pinned SmolLM3-3B-Base before evaluation.
Full 1319 test examples, zero-shot, original extractive_match: 0.4086429112964367 (stderr 0.013540639733342422).
Free Google Colab T4, no paid HF Jobs; cost 0.
lighteval 0.11.0, vLLM 0.10.1.1, Transformers 4.57.1, Python 3.12.
Dataset-address correction… See the full description on the dataset page: https://huggingface.co/datasets/bestdive/details_bestdive__SmolLM3-3B-SFT-Free-Course.RULER-65536-SmolLM3-SFTRULER-262144-SmolLM3-11T-32k-v1-remote-codesmollm3-stack-v2-C-Sharptis-subset-datasets-SmolLM3-3B-Basesmollm3-stack-v2-HTML
