datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jupyter-scripts-smollm3
The Stack v2 Jupyter Notebooks as Scripts
This dataset contains script representations of the Jupyter notebooks in
The Stack v2. It was
created from the materialized Jupyter_Notebook split in
jordangong/the-stack-v2-smollm3.
The output schema follows the Jupyter-script schema used by
bigcode/starcoderdata,
but this release is not deduplicated, PII-filtered, or otherwise equivalent
to StarCoderData's filtered split.
Relationship to the SmolLM3 training mix
This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.pretraining-pretokenized-smollm3
SmolLM3 Pretokenized Pretraining Sources
Datatrove/Nanotron tokenized-byte versions of three public pretraining sources:
fineweb-edu-10bt: HuggingFaceFW/fineweb-edu, sample/10BT
finemath-4plus: HuggingFaceTB/finemath, finemath-4plus
stack-edu-python: HuggingFaceTB/stack-edu, Python, with content retrieved from the public Software Heritage S3 bucket using the upstream dataset-card procedure
All subsets use HuggingFaceTB/SmolLM3-3B at revision… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/pretraining-pretokenized-smollm3.ai_human_smollm360m_logits
AI/Human logits — SmolLM-360M
Next-token distributions from HuggingFaceTB/SmolLM-360M over
AI_Human_Dataset.csv from Aanimated/telescope_datasets.
At each position, only tokens with softmax probability >= the cutoff are
kept (the argmax is always retained), sorted by descending probability.
Columns
column
description
sample_index
index of the source row
input_ids
SmolLM token ids for the sample
num_tokens
sequence length after truncation
token_ids… See the full description on the dataset page: https://huggingface.co/datasets/Joshfcooper/ai_human_smollm360m_logits.smollm3-tracesportuguese-eval-logs-olmo2-smollm3
Evaluation Logs on Portuguese Benchmarks for OLMo-2 and SmolLM3
These logs contain benchmark results across a suite of Portuguese-language tasks. The data consists of recordings of the performance of various 3 different models at different checkpoints throughout their pretraining runs:
SmolLM3
OLMo-2-0425-1B
OLMo-2-1124-7B
Splits
Each split (smollm3_3b, olmo2_1b, olmo2_7b) contains rows for model checkpoints and columns for benchmark scores (e.g., ASSIN2 RTE, ENEM… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-eval-logs-olmo2-smollm3.RULER-131072-SmolLM3-11T-32k-v1-remote-codeRULER-262144-SmolLM3-11T-32k-v1-remote-codeRULER-65536-SmolLM3-11T-32k-v1-remote-codeRULER-4096-SmolLM3-11T-32k-v1-remote-codedetails_pmakiela__SmolLM3-3B-dpo-v0_1_private
Dataset Card for Evaluation run of pmakiela/SmolLM3-3B-dpo-v0_1
Dataset automatically created during the evaluation run of model pmakiela/SmolLM3-3B-dpo-v0_1.
The dataset is composed of 4 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/pmakiela/details_pmakiela__SmolLM3-3B-dpo-v0_1_private.RULER-32768-SmolLM3-11T-32k-v1-remote-codedetails_quablab__SmolLM3-Custom-SFT_private
Dataset Card for Evaluation run of quablab/SmolLM3-Custom-SFT
Dataset automatically created during the evaluation run of model quablab/SmolLM3-Custom-SFT.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/quablab/details_quablab__SmolLM3-Custom-SFT_private.details_onethreedlee__SmolLM3-Custom-SFT_private
Dataset Card for Evaluation run of onethreedlee/SmolLM3-Custom-SFT
Dataset automatically created during the evaluation run of model onethreedlee/SmolLM3-Custom-SFT.
The dataset is composed of 2 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/onethreedlee/details_onethreedlee__SmolLM3-Custom-SFT_private.details_HuggingFaceTB__SmolLM3-3B_private
Dataset Card for Evaluation run of HuggingFaceTB/SmolLM3-3B
Dataset automatically created during the evaluation run of model HuggingFaceTB/SmolLM3-3B.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/andregustavo/details_HuggingFaceTB__SmolLM3-3B_private.details_iamgroot42__smollm3-Custom-SFT
Dataset Card for Evaluation run of iamgroot42/smollm3-Custom-SFT
Dataset automatically created during the evaluation run of model iamgroot42/smollm3-Custom-SFT.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/iamgroot42/details_iamgroot42__smollm3-Custom-SFT.details_jweston__SmolLM3-Custom-SFT_private
Dataset Card for Evaluation run of jweston/SmolLM3-Custom-SFT
Dataset automatically created during the evaluation run of model jweston/SmolLM3-Custom-SFT.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/jweston/details_jweston__SmolLM3-Custom-SFT_private.details-smollm3-3b-sft-finetuned-jobs-v2
Dataset Card for Evaluation run of marcelovidigal/smollm3-3b-sft-finetuned-jobs-v2
Dataset automatically created during the evaluation run of model marcelovidigal/smollm3-3b-sft-finetuned-jobs-v2.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/marcelovidigal/details-smollm3-3b-sft-finetuned-jobs-v2.RULER-16384-SmolLM3-11T-32k-v1-remote-codedetails_lukmanaj__smollm3-sft-colab-merged_private
Dataset Card for Evaluation run of lukmanaj/smollm3-sft-colab-merged
Dataset automatically created during the evaluation run of model lukmanaj/smollm3-sft-colab-merged.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/lukmanaj/details_lukmanaj__smollm3-sft-colab-merged_private.RULER-8192-SmolLM3-11T-32k-v1-remote-codedetails_HuggingFaceTB__SmolLM3-3B_private
Dataset Card for Evaluation run of HuggingFaceTB/SmolLM3-3B
Dataset automatically created during the evaluation run of model HuggingFaceTB/SmolLM3-3B.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/smol-course/details_HuggingFaceTB__SmolLM3-3B_private.details_pmakiela__SmolLM3-3B-SFT-v1_4_private
Dataset Card for Evaluation run of pmakiela/SmolLM3-3B-SFT-v1_4
Dataset automatically created during the evaluation run of model pmakiela/SmolLM3-3B-SFT-v1_4.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/pmakiela/details_pmakiela__SmolLM3-3B-SFT-v1_4_private.details_quablab__smollm3-dpo-aligned_private
Dataset Card for Evaluation run of quablab/smollm3-dpo-aligned
Dataset automatically created during the evaluation run of model quablab/smollm3-dpo-aligned.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/quablab/details_quablab__smollm3-dpo-aligned_private.smollm3-3b-base-blind-spots
Blind Spots of SmolLM3-3B-Base
This dataset contains 15 diverse examples where the base language model HuggingFaceTB/SmolLM3-3B-Base produces incorrect, repetitive, or off‑task outputs. It was created as part of the Fatima Fellowship technical challenge.
Model
Name: HuggingFaceTB/SmolLM3-3B-Base
Type: Decoder‑only transformer (base model, not instruction‑tuned)
Parameters: 3B
Link: https://huggingface.co/HuggingFaceTB/SmolLM3-3B-Base
Methodology
I loaded the… See the full description on the dataset page: https://huggingface.co/datasets/AyshSaleem/smollm3-3b-base-blind-spots.details_emharsha1812__smollm3-lora-full-merged_private
Dataset Card for Evaluation run of emharsha1812/smollm3-lora-full-merged
Dataset automatically created during the evaluation run of model emharsha1812/smollm3-lora-full-merged.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/emharsha1812/details_emharsha1812__smollm3-lora-full-merged_private.details_pmakiela__SmolLM3-3B-SFT-v1_2_private
Dataset Card for Evaluation run of pmakiela/SmolLM3-3B-SFT-v1_2
Dataset automatically created during the evaluation run of model pmakiela/SmolLM3-3B-SFT-v1_2.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/pmakiela/details_pmakiela__SmolLM3-3B-SFT-v1_2_private.details_kshitijthakkar__SmolLM3-Custom-SFT_private
Dataset Card for Evaluation run of kshitijthakkar/SmolLM3-Custom-SFT
Dataset automatically created during the evaluation run of model kshitijthakkar/SmolLM3-Custom-SFT.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/details_kshitijthakkar__SmolLM3-Custom-SFT_private.smollm3-3b-dmmath-tracesSmolLM3-3B-customerservice-LLM-as-a-judge-datasmollm3-base-blindspots
SmolLM3-3B-Base Blind Spots Evaluation Dataset
Dataset Summary
This dataset documents 10 diverse failure cases discovered while evaluating
HuggingFaceTB/SmolLM3-3B-Base,
a 3-billion parameter decoder-only base language model released by Hugging Face in July 2025.
The evaluation was conducted as part of the Fatima Fellowship technical challenge on Blind Spots of Frontier Models.
Model Tested
Model: HuggingFaceTB/SmolLM3-3B-Base
Parameters: 3 billion… See the full description on the dataset page: https://huggingface.co/datasets/habibahabchi/smollm3-base-blindspots.
