datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jupyter-scripts-smollm3
The Stack v2 Jupyter Notebooks as Scripts
This dataset contains script representations of the Jupyter notebooks in
The Stack v2. It was
created from the materialized Jupyter_Notebook split in
jordangong/the-stack-v2-smollm3.
The output schema follows the Jupyter-script schema used by
bigcode/starcoderdata,
but this release is not deduplicated, PII-filtered, or otherwise equivalent
to StarCoderData's filtered split.
Relationship to the SmolLM3 training mix
This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.matilda-smollm-mix-15b-gpt2
matilda-smollm-mix-15B-gpt2
15 B GPT-2-BPE tokens drawn from a 5:1 token-balanced mix of
HuggingFaceTB/smollm-corpus:
Source
Share
Tokens
fineweb-edu-dedup
83.33 %
12.50 B
cosmopedia-v2
16.67 %
2.50 B
Total: 15,000,349,569 tokens across 151 shards (shard_*.bin, uint16,
100 M tokens per shard).
The full SmolLM recipe is 75 / 15 / 10 fineweb-edu / cosmopedia-v2 / python-edu.
python-edu was dropped because the HuggingFaceTB/smollm-corpus subset
ships only blob_id… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/matilda-smollm-mix-15b-gpt2.smollm3-base-blindspots
SmolLM3-3B-Base Blind Spots Evaluation Dataset
Dataset Summary
This dataset documents 10 diverse failure cases discovered while evaluating
HuggingFaceTB/SmolLM3-3B-Base,
a 3-billion parameter decoder-only base language model released by Hugging Face in July 2025.
The evaluation was conducted as part of the Fatima Fellowship technical challenge on Blind Spots of Frontier Models.
Model Tested
Model: HuggingFaceTB/SmolLM3-3B-Base
Parameters: 3 billion… See the full description on the dataset page: https://huggingface.co/datasets/habibahabchi/smollm3-base-blindspots.smollm-corpus-python
smollm-corpus - python
A version of the python-edu subset with the text added
