datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-stack-v2-smollm3
The Stack v2 — materialized source code
Upstream dataset:
bigcode/the-stack-v2
Exact upstream commit:
e565caa3a78c2423bd374333a472b049eb090e47
Primary source-content endpoint:
https://softwareheritage.s3.amazonaws.com/content/{blob_id}
Configurations
TypeScript
Swift
Ruby
Rust
Go
Shell
Jupyter_Notebook
HTML
Python
Java
JavaScript
C
C++
C-Sharp
PHP
SQL
Markdown
Added columns
content: decoded source content
download_error: null on successful… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/the-stack-v2-smollm3.jupyter-scripts-smollm3
The Stack v2 Jupyter Notebooks as Scripts
This dataset contains script representations of the Jupyter notebooks in
The Stack v2. It was
created from the materialized Jupyter_Notebook split in
jordangong/the-stack-v2-smollm3.
The output schema follows the Jupyter-script schema used by
bigcode/starcoderdata,
but this release is not deduplicated, PII-filtered, or otherwise equivalent
to StarCoderData's filtered split.
Relationship to the SmolLM3 training mix
This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.c4-rewritten-14b-retok-smollm360mdclm-14b-c4-rewritten-14b-retok-smollm360mdclm-6.7b-c4-rewritten-6.7b-retok-smollm360mc4-rewritten-6.7b-retok-smollm360mpretraining-pretokenized-smollm3
SmolLM3 Pretokenized Pretraining Sources
Datatrove/Nanotron tokenized-byte versions of three public pretraining sources:
fineweb-edu-10bt: HuggingFaceFW/fineweb-edu, sample/10BT
finemath-4plus: HuggingFaceTB/finemath, finemath-4plus
stack-edu-python: HuggingFaceTB/stack-edu, Python, with content retrieved from the public Software Heritage S3 bucket using the upstream dataset-card procedure
All subsets use HuggingFaceTB/SmolLM3-3B at revision… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/pretraining-pretokenized-smollm3.smollm3-tracesportuguese-eval-logs-olmo2-smollm3
Evaluation Logs on Portuguese Benchmarks for OLMo-2 and SmolLM3
These logs contain benchmark results across a suite of Portuguese-language tasks. The data consists of recordings of the performance of various 3 different models at different checkpoints throughout their pretraining runs:
SmolLM3
OLMo-2-0425-1B
OLMo-2-1124-7B
Splits
Each split (smollm3_3b, olmo2_1b, olmo2_7b) contains rows for model checkpoints and columns for benchmark scores (e.g., ASSIN2 RTE, ENEM… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-eval-logs-olmo2-smollm3.RULER-131072-SmolLM3-11T-32k-v1-remote-codedetails_bestdive__SmolLM3-3B-SFT-Free-Course
Smol course SFT evaluation - Kay Zheng
Actual full GSM8K test evaluation of bestdive/SmolLM3-3B-SFT-Free-Course, adapter revision 0484e028b494d605a267050a949c9266edadd16b, merged with pinned SmolLM3-3B-Base before evaluation.
Full 1319 test examples, zero-shot, original extractive_match: 0.4086429112964367 (stderr 0.013540639733342422).
Free Google Colab T4, no paid HF Jobs; cost 0.
lighteval 0.11.0, vLLM 0.10.1.1, Transformers 4.57.1, Python 3.12.
Dataset-address correction… See the full description on the dataset page: https://huggingface.co/datasets/bestdive/details_bestdive__SmolLM3-3B-SFT-Free-Course.RULER-262144-SmolLM3-11T-32k-v1-remote-codetis-subset-datasets-SmolLM3-3B-BaseRULER-65536-SmolLM3-11T-32k-v1-remote-codeRULER-4096-SmolLM3-11T-32k-v1-remote-codedetails_pmakiela__SmolLM3-3B-dpo-v0_1_private
Dataset Card for Evaluation run of pmakiela/SmolLM3-3B-dpo-v0_1
Dataset automatically created during the evaluation run of model pmakiela/SmolLM3-3B-dpo-v0_1.
The dataset is composed of 4 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/pmakiela/details_pmakiela__SmolLM3-3B-dpo-v0_1_private.details_serverdaun__smollm3-dpo
Dataset Card for Evaluation run of serverdaun/smollm3-dpo
Dataset automatically created during the evaluation run of model serverdaun/smollm3-dpo.
The dataset is composed of 3 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/iamtokarev/details_serverdaun__smollm3-dpo.tis-dolci-subset-datasets-SmolLM3-3B-Basedetails_FrancescoArno94__SmolLM3-3B-math_private
Dataset Card for Evaluation run of FrancescoArno94/SmolLM3-3B-math
Dataset automatically created during the evaluation run of model FrancescoArno94/SmolLM3-3B-math.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/FrancescoArno94/details_FrancescoArno94__SmolLM3-3B-math_private.jackhhao_jailbreak_classification_smolLM3_qwen3-4BsmolLM3_french_data
Description
Hugging Face's SmolLM3 was introduced along with its training dataset: smoltalk2.This dataset includes the smoltalk_multilingual8_Qwen3_32B_think split in the SFT subset, which contains multilingual data, including French.However, there is no column available to filter this data by language.One option would be to apply a language detection model, but to avoid any errors, we requested the original French dataset, before it was mixed with other languages, to the… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/smolLM3_french_data.MMLU-Pro_SmolLM3-3B_test
MMLU-Pro_SmolLM3-3B_test
RULER-32768-SmolLM3-11T-32k-v1-remote-codedetails_yanzhiqiang__SmolLM3-Custom-SFT_private
Dataset Card for Evaluation run of yanzhiqiang/SmolLM3-Custom-SFT
Dataset automatically created during the evaluation run of model yanzhiqiang/SmolLM3-Custom-SFT.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/yanzhiqiang/details_yanzhiqiang__SmolLM3-Custom-SFT_private.details_Abdullahpk1982__smollm3-dpo-aligned_private
Dataset Card for Evaluation run of Abdullahpk1982/smollm3-dpo-aligned
Dataset automatically created during the evaluation run of model Abdullahpk1982/smollm3-dpo-aligned.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/Abdullahpk1982/details_Abdullahpk1982__smollm3-dpo-aligned_private.details_yanzhiqiang__smollm3-dpo-aligned_private
Dataset Card for Evaluation run of yanzhiqiang/smollm3-dpo-aligned
Dataset automatically created during the evaluation run of model yanzhiqiang/smollm3-dpo-aligned.
The dataset is composed of 3 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/yanzhiqiang/details_yanzhiqiang__smollm3-dpo-aligned_private.details_quablab__SmolLM3-Custom-SFT_private
Dataset Card for Evaluation run of quablab/SmolLM3-Custom-SFT
Dataset automatically created during the evaluation run of model quablab/SmolLM3-Custom-SFT.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/quablab/details_quablab__SmolLM3-Custom-SFT_private.details_marinarosa_smollm3-finetuned_lighteval_gsm8k
Dataset Card for Evaluation run of marinarosa/smollm3-finetuned
Dataset automatically created during the evaluation run of model marinarosa/smollm3-finetuned.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/marinarosa/details_marinarosa_smollm3-finetuned_lighteval_gsm8k.details_sidgenai__smollm3-sft_private
Dataset Card for Evaluation run of sidgenai/smollm3-sft
Dataset automatically created during the evaluation run of model sidgenai/smollm3-sft.
The dataset is composed of 1 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/sidgenai/details_sidgenai__smollm3-sft_private.jailbreak_classification_smolLM3
