CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceTB /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus.tabular100M<n<1B486 likes54k downloads2y agoHugging Face02enguyen /smollm-chunked FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity. Full Documentation For complete usage instructions, installation guide, and tutorial, please refer to: Main Tutorial README Data Distribution Due to Hugging Face storage quota… See the full description on the dataset page: https://huggingface.co/datasets/enguyen/smollm-chunked.tabulartext-retrieval100M<n<1B1 likes35k downloads7mo agoHugging Face03chengjunyan1 /smollm-12.5-corpus SmolLM-1/8-Corpus Around 1/8 upper-quality subset of SmolLM Corpus for training Chinchilla-optimal GPT-2 scale (sub 1.5B) models, which is a good scale for verifying a model architecture under the scaling laws. Firstly filtered samples with int_score >=4 from FineWeb-edu-dedup, then keep the training mixture with the same distribution from SmolLM. In which FineWeb-Edu-dedup occupies around 70% of the corpus. Then sample other dataset based on the mixture ratios respectively. For… See the full description on the dataset page: https://huggingface.co/datasets/chengjunyan1/smollm-12.5-corpus.tabular100M<n<1B7 likes1.8k downloads2y agoHugging Face04aklein4 /seq2seq-mixed-pretraining-SmolLM2tabular100M<n<1B1 likes1.6k downloads8mo agoHugging Face05jordangong /jupyter-scripts-smollm3 The Stack v2 Jupyter Notebooks as Scripts This dataset contains script representations of the Jupyter notebooks in The Stack v2. It was created from the materialized Jupyter_Notebook split in jordangong/the-stack-v2-smollm3. The output schema follows the Jupyter-script schema used by bigcode/starcoderdata, but this release is not deduplicated, PII-filtered, or otherwise equivalent to StarCoderData's filtered split. Relationship to the SmolLM3 training mix This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.tabulartext-generation1M<n<10M0 likes1.1k downloads15d agoHugging Face06chengjunyan1 /smollm-10tabular10M<n<100M0 likes839 downloads2y agoHugging Face07styal /smollm-corpus-3.5M A very small version of smollm-corpus (+finemath-4plus) for experimenting llm pre-training. cosmopedia-v2 (1M rows) fineweb-edu-dedup (1M rows) python-edu (0.5M rows) finemath-4plus (1M rows) tabular1M<n<10M2 likes428 downloads2y agoHugging Face08oieieio /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/oieieio/smollm-corpus.tabular100M<n<1B0 likes218 downloads1y agoHugging Face09juiceb0xc0de /smollm2-135m-instruct-SAE Layer EV Mean L0 Recon Loss Dead % 0 0.9480 48.74 0.2074 0.0 1 0.9599 43.65 0.3298 0.0 2 0.9631 46.81 0.5021 0.0 3 0.9508 46.56 0.7462 0.0 4 0.9463 46.23 0.8936 0.0 5 0.9350 47.57 1.1605 0.0 6 0.9306 48.44 1.3838 0.0 7 0.9318 49.51 1.5446 0.0 8 0.9432 46.52 1.6598 0.0 9 0.9373 47.15 2.0706 0.0 10 0.9348 45.53 2.2983 0.0 11 0.9905 48.58 5.8113 0.0 12 0.9901 48.42 6.1039 0.0 13 0.9891 46.15 6.9692 0.0 14 0.9884 44.76 7.1844 0.0 15 0.9863 47.63 8.6521 0.0… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/smollm2-135m-instruct-SAE.tabularfeature-extractionn<1K0 likes180 downloads3mo agoHugging Face10aklein4 /Bitext-SmolLM2-1024-natural-instructions-formattabularn<1K0 likes161 downloads3mo agoHugging Face11Joshfcooper /ai_human_smollm360m_logits AI/Human logits — SmolLM-360M Next-token distributions from HuggingFaceTB/SmolLM-360M over AI_Human_Dataset.csv from Aanimated/telescope_datasets. At each position, only tokens with softmax probability >= the cutoff are kept (the argmax is always retained), sorted by descending probability. Columns column description sample_index index of the source row input_ids SmolLM token ids for the sample num_tokens sequence length after truncation token_ids… See the full description on the dataset page: https://huggingface.co/datasets/Joshfcooper/ai_human_smollm360m_logits.tabulartext-classification100K<n<1M0 likes159 downloads2mo agoHugging Face12akseljoonas /smollm3-tracestabularn<1K0 likes151 downloads1y agoHugging Face13Polygl0t /portuguese-eval-logs-olmo2-smollm3 Evaluation Logs on Portuguese Benchmarks for OLMo-2 and SmolLM3 These logs contain benchmark results across a suite of Portuguese-language tasks. The data consists of recordings of the performance of various 3 different models at different checkpoints throughout their pretraining runs: SmolLM3 OLMo-2-0425-1B OLMo-2-1124-7B Splits Each split (smollm3_3b, olmo2_1b, olmo2_7b) contains rows for model checkpoints and columns for benchmark scores (e.g., ASSIN2 RTE, ENEM… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-eval-logs-olmo2-smollm3.imagen<1K0 likes149 downloads7mo agoHugging Face14ZhuofengLi /pretraining-pretokenized-smollm3 SmolLM3 Pretokenized Pretraining Sources Datatrove/Nanotron tokenized-byte versions of three public pretraining sources: fineweb-edu-10bt: HuggingFaceFW/fineweb-edu, sample/10BT finemath-4plus: HuggingFaceTB/finemath, finemath-4plus stack-edu-python: HuggingFaceTB/stack-edu, Python, with content retrieved from the public Software Heritage S3 bucket using the upstream dataset-card procedure All subsets use HuggingFaceTB/SmolLM3-3B at revision… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/pretraining-pretokenized-smollm3.tabularn<1K0 likes149 downloads2mo agoHugging Face15lighteval /RULER-262144-SmolLM3-11T-32k-v1-remote-codetabular1K<n<10K0 likes97 downloads1y agoHugging Face16lighteval /RULER-131072-SmolLM3-11T-32k-v1-remote-codetabular1K<n<10K0 likes93 downloads1y agoHugging Face17styal /very-smollm-corpus-0.5Mtabular100K<n<1M0 likes68 downloads2y agoHugging Face18malaiwah /qfs-smollm2-135m-wikitext2-native-v1 HF workflow d3dc69602aeb981f06bd9f4c726937f9 A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-native-bf16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-native-v1.tabularn<1K0 likes67 downloads17d agoHugging Face19open-llm-leaderboard /HuggingFaceTB__SmolLM-1.7B-Instruct-detailsgated Dataset Card for Evaluation run of HuggingFaceTB/SmolLM-1.7B-Instruct Dataset automatically created during the evaluation run of model HuggingFaceTB/SmolLM-1.7B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceTB__SmolLM-1.7B-Instruct-details.tabular10K<n<100K0 likes59 downloads2y agoHugging Face20hhoenjet /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/hhoenjet/smollm-corpus.tabular100M<n<1B0 likes56 downloads7mo agoHugging Face21malaiwah /qfs-smollm2-135m-wikitext2-gptq-g32-v1 HF workflow 32c6ab05b0ceab1cecdceda838846388 A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-gptq-int4-g32. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-gptq-g32-v1.tabularn<1K0 likes55 downloads17d agoHugging Face22lighteval /RULER-65536-SmolLM3-11T-32k-v1-remote-codetabular1K<n<10K0 likes51 downloads1y agoHugging Face23malaiwah /qfs-smollm2-135m-wikitext2-gptq-g64-v1 HF workflow 73f0a12a901c7368794a3a886f55b675 A quant fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from malaiwah/SmolLM2-135M-QFS-gptq-int4-g64. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qfs-smollm2-135m-wikitext2-gptq-g64-v1.tabularn<1K0 likes49 downloads17d agoHugging Face24lighteval /RULER-4096-SmolLM3-11T-32k-v1-remote-codetabular1K<n<10K0 likes47 downloads1y agoHugging Face25Mgmgrand420 /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/Mgmgrand420/smollm-corpus.tabular100M<n<1B0 likes43 downloads8mo agoHugging Face26open-llm-leaderboard /meditsolutions__SmolLM2-MedIT-Upscale-2B-detailsgated Dataset Card for Evaluation run of meditsolutions/SmolLM2-MedIT-Upscale-2B Dataset automatically created during the evaluation run of model meditsolutions/SmolLM2-MedIT-Upscale-2B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meditsolutions__SmolLM2-MedIT-Upscale-2B-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face27open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_1000k_fineweb-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_1000k_fineweb Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_1000k_fineweb The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_1000k_fineweb-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face28open-llm-leaderboard /FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_selected-detailsgated Dataset Card for Evaluation run of FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_selected Dataset automatically created during the evaluation run of model FlofloB/smollm2-135M_pretrained_800k_fineweb_uncovai_selected The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__smollm2-135M_pretrained_800k_fineweb_uncovai_selected-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face29pmakiela /details_pmakiela__SmolLM3-3B-dpo-v0_1_private Dataset Card for Evaluation run of pmakiela/SmolLM3-3B-dpo-v0_1 Dataset automatically created during the evaluation run of model pmakiela/SmolLM3-3B-dpo-v0_1. The dataset is composed of 4 configuration, each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/pmakiela/details_pmakiela__SmolLM3-3B-dpo-v0_1_private.tabular10K<n<100K0 likes41 downloads1y agoHugging Face30aklein4 /compilation-SmolLM2tabular10M<n<100M0 likes41 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.