datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SmolLM2-135M-10BThis dataset is sampled from the SmolLM2 Corpus described in https://arxiv.org/abs/2502.02737. Specifically, we sampled from
the SmolLM2-135M pretraining data, a 2T token mixture consisting of four complete high quality datasets, and selected portions of
DCLM-Edu and FineWeb-Edu sampled at a 6:4 ratio.
This sample is intended to enable fast downloading and training of sparsify models.
FineMath: 34B tokens
Stack-Edu: 125B tokens
InfiMM-WebMath: 40B tokens
Cosmopedia V2: 30B tokens… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/SmolLM2-135M-10B.seq2seq-mixed-pretraining-SmolLM2math-rlvr-mini-smollm2-0.4b-v2SmolLM2-1.7B-stage-4-20BSmolLM2-1.7B-stage-4-100Bbooks3-SmolLM2-sortedbooks3-SmolLM2details_HuggingFaceTB__SmolLM2-1.7B-Instruct
Dataset Card for Evaluation run of HuggingFaceTB/SmolLM2-1.7B-Instruct
Dataset automatically created during the evaluation run of model HuggingFaceTB/SmolLM2-1.7B-Instruct.
The dataset is composed of 7 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 12 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_HuggingFaceTB__SmolLM2-1.7B-Instruct.SmolLM2-135M-20BSmolLM2-135M-100Bsmollm2-135m-instruct-SAE
Layer
EV
Mean L0
Recon Loss
Dead %
0
0.9480
48.74
0.2074
0.0
1
0.9599
43.65
0.3298
0.0
2
0.9631
46.81
0.5021
0.0
3
0.9508
46.56
0.7462
0.0
4
0.9463
46.23
0.8936
0.0
5
0.9350
47.57
1.1605
0.0
6
0.9306
48.44
1.3838
0.0
7
0.9318
49.51
1.5446
0.0
8
0.9432
46.52
1.6598
0.0
9
0.9373
47.15
2.0706
0.0
10
0.9348
45.53
2.2983
0.0
11
0.9905
48.58
5.8113
0.0
12
0.9901
48.42
6.1039
0.0
13
0.9891
46.15
6.9692
0.0
14
0.9884
44.76
7.1844
0.0
15
0.9863
47.63
8.6521
0.0… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/smollm2-135m-instruct-SAE.Bitext-SmolLM2-1024-natural-instructions-formatsingle-turn-compilation-SmolLM2-1024SmolLM2-1.7B-stage-4-10Bmath-rlvr-mini-sa-smollm2-0.1b-v1math-rlvr-mini-sa-smollm2-0.1b-v0math-rlvr-mini-sa-smollm2-0.1b-v2qfs-smollm2-135m-wikitext2-campaign-v1
SmolLM2-135M QFS calibration and evaluation campaign
A small stored-weight fidelity study, not a broad model-quality benchmark.
Evaluation: 16 complete WikiText2 raw test articles, one 256-token window each, 4080 prediction positions.
Calibration: 32 disjoint train articles, 256 tokens each, 8192 calibration tokens.
Complete article title, normalized content and exact 13-token-ngram separation were checked. Validation is unused. Pretraining overlap remains unknown.
Original… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-campaign-v1.math-rlvr-mini-sa-smollm2-0.4b-v2math-rlvr-mini-sa-smollm2-0.4b-v0math-rlvr-mini-smollm2-1.7b-v2qfs-smollm2-135m-wikitext2-native-v1
HF workflow d3dc69602aeb981f06bd9f4c726937f9
A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-native-bf16.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-native-v1.ai-vs-human-HuggingFaceTB-SmolLM2-1.7B-Instruct
AI vs Human dataset on the CNN Daily mails
Dataset Description
This dataset showcases pairs of truncated articles and their respective completions, crafted either by humans or an AI language model.
Each article was randomly truncated between 25% and 50% of its length.
The language model was then tasked with generating a completion that mirrored the characters count of the original human-written continuation.
Data Fields
'human': The original human-authored… See the full description on the dataset page: https://huggingface.co/datasets/zcamz/ai-vs-human-HuggingFaceTB-SmolLM2-1.7B-Instruct.OpenR1-Math-220k-200M_SmolLM2-1.7Bqfs-smollm2-135m-wikitext2-gptq-g32-v1
HF workflow 32c6ab05b0ceab1cecdceda838846388
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-gptq-int4-g32.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-gptq-g32-v1.compilation-SmolLM2qfs-smollm2-135m-wikitext2-gptq-g64-v1
HF workflow 73f0a12a901c7368794a3a886f55b675
A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-gptq-int4-g64.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-gptq-g64-v1.math-rlvr-mini-sa-smollm2-0.4b-v1math-rlvr-mini-sa-smollm2-0.4b-v3smollm2-tool-calling-sft-data
