smollm3
Datasets
All datasets matching “smollm3”the-stack-v2-smollm3
The Stack v2 — materialized source code
Upstream dataset:
bigcode/the-stack-v2
Exact upstream commit:
e565caa3a78c2423bd374333a472b049eb090e47
Primary source-content endpoint:
https://softwareheritage.s3.amazonaws.com/content/{blob_id}
Configurations
TypeScript
Swift
Ruby
Rust
Go
Shell
Jupyter_Notebook
HTML
Python
Java
JavaScript
C
C++
C-Sharp
PHP
SQL
Markdown
Added columns
content: decoded source content
download_error: null on successful… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/the-stack-v2-smollm3.jupyter-scripts-smollm3
The Stack v2 Jupyter Notebooks as Scripts
This dataset contains script representations of the Jupyter notebooks in
The Stack v2. It was
created from the materialized Jupyter_Notebook split in
jordangong/the-stack-v2-smollm3.
The output schema follows the Jupyter-script schema used by
bigcode/starcoderdata,
but this release is not deduplicated, PII-filtered, or otherwise equivalent
to StarCoderData's filtered split.
Relationship to the SmolLM3 training mix
This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.smollm3-configs
SmolLM3 Training Configs
[IMPORTANT NOTE]: for the latest configs go to this repo: https://github.com/huggingface/smollm/tree/main/text/pretraining/smollm3
Here you can find the training configs for SmoLLM3-3B-Base using nanotron with exact training details and data mixtures.
The model was trained on 11.2T tokens in 3 stages on 4k context:
stage 1 config
stage 2 config
stage 3 config
And then we trained on an additional 2 stages to extend the contetx length to 64k:
stage 4… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/smollm3-configs.c4-rewritten-14b-retok-smollm360msmollm3-baseline_v3dclm-14b-c4-rewritten-14b-retok-smollm360m
