CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jordangong /the-stack-v2-smollm3 The Stack v2 — materialized source code Upstream dataset: bigcode/the-stack-v2 Exact upstream commit: e565caa3a78c2423bd374333a472b049eb090e47 Primary source-content endpoint: https://softwareheritage.s3.amazonaws.com/content/{blob_id} Configurations TypeScript Swift Ruby Rust Go Shell Jupyter_Notebook HTML Python Java JavaScript C C++ C-Sharp PHP SQL Markdown Added columns content: decoded source content download_error: null on successful… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/the-stack-v2-smollm3.texttext-generation1B<n<10B1 likes14k downloads14d agoHugging Face02jordangong /jupyter-scripts-smollm3 The Stack v2 Jupyter Notebooks as Scripts This dataset contains script representations of the Jupyter notebooks in The Stack v2. It was created from the materialized Jupyter_Notebook split in jordangong/the-stack-v2-smollm3. The output schema follows the Jupyter-script schema used by bigcode/starcoderdata, but this release is not deduplicated, PII-filtered, or otherwise equivalent to StarCoderData's filtered split. Relationship to the SmolLM3 training mix This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.tabulartext-generation1M<n<10M0 likes1.1k downloads15d agoHugging Face03HuggingFaceTB /smollm3-configs SmolLM3 Training Configs [IMPORTANT NOTE]: for the latest configs go to this repo: https://github.com/huggingface/smollm/tree/main/text/pretraining/smollm3 Here you can find the training configs for SmoLLM3-3B-Base using nanotron with exact training details and data mixtures. The model was trained on 11.2T tokens in 3 stages on 4k context: stage 1 config stage 2 config stage 3 config And then we trained on an additional 2 stages to extend the contetx length to 64k: stage 4… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/smollm3-configs.8 likes633 downloads1y agoHugging Face04ardauzunoglu /c4-rewritten-14b-retok-smollm360mtext10M<n<100M0 likes565 downloads2mo agoHugging Face05verify-ppt /smollm3-baseline_v30 likes512 downloads3mo agoHugging Face06ardauzunoglu /dclm-14b-c4-rewritten-14b-retok-smollm360mtext10M<n<100M0 likes496 downloads2mo agoHugging Face07ardauzunoglu /dclm-6.7b-c4-rewritten-6.7b-retok-smollm360mtext10M<n<100M0 likes460 downloads2mo agoHugging Face08HuggingFaceTB /smollm3-blueprintHere you can find the SmolLM3 Engineering Blueprint documentn<1K9 likes278 downloads1y agoHugging Face09ardauzunoglu /c4-rewritten-6.7b-retok-smollm360mtext1M<n<10M0 likes262 downloads2mo agoHugging Face10ZhuofengLi /pretraining-pretokenized-smollm3 SmolLM3 Pretokenized Pretraining Sources Datatrove/Nanotron tokenized-byte versions of three public pretraining sources: fineweb-edu-10bt: HuggingFaceFW/fineweb-edu, sample/10BT finemath-4plus: HuggingFaceTB/finemath, finemath-4plus stack-edu-python: HuggingFaceTB/stack-edu, Python, with content retrieved from the public Software Heritage S3 bucket using the upstream dataset-card procedure All subsets use HuggingFaceTB/SmolLM3-3B at revision… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/pretraining-pretokenized-smollm3.tabularn<1K0 likes161 downloads2mo agoHugging Face11verify-ppt /smollm3-wiki0 likes159 downloads6mo agoHugging Face12reasoning-cues /rollouts-smollm3-rl rollouts-smollm3-rl — evaluation rollouts (boxed prompt) Model: reasoning-cues/smollm3-rl @ 97bda61 — a GRPO LoRA adapter (r=64, alpha=128, all attention and MLP projections) on HuggingFaceTB/SmolLM3-3B-Base @ d78a42f, merged with peft merge_and_unload (bf16) and served with the base model's tokenizer files (the adapter repo's tokenizer_config.json declares TokenizersBackend; its tokenizer.json is byte-identical to the base's). Protocol: the paper's released-model grid — 32… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-smollm3-rl.0 likes159 downloads7d agoHugging Face13Joshfcooper /ai_human_smollm360m_logits AI/Human logits — SmolLM-360M Next-token distributions from HuggingFaceTB/SmolLM-360M over AI_Human_Dataset.csv from Aanimated/telescope_datasets. At each position, only tokens with softmax probability >= the cutoff are kept (the argmax is always retained), sorted by descending probability. Columns column description sample_index index of the source row input_ids SmolLM token ids for the sample num_tokens sequence length after truncation token_ids… See the full description on the dataset page: https://huggingface.co/datasets/Joshfcooper/ai_human_smollm360m_logits.tabulartext-classification100K<n<1M0 likes157 downloads2mo agoHugging Face14verify-ppt /smollm3-baseline_v20 likes156 downloads4mo agoHugging Face15akseljoonas /smollm3-tracestabularn<1K0 likes150 downloads1y agoHugging Face16verify-ppt /smollm3-infiwebmath0 likes133 downloads6mo agoHugging Face17verify-ppt /smollm3-github-issues0 likes131 downloads6mo agoHugging Face18verify-ppt /smollm3-stackexchange0 likes129 downloads6mo agoHugging Face19Polygl0t /portuguese-eval-logs-olmo2-smollm3 Evaluation Logs on Portuguese Benchmarks for OLMo-2 and SmolLM3 These logs contain benchmark results across a suite of Portuguese-language tasks. The data consists of recordings of the performance of various 3 different models at different checkpoints throughout their pretraining runs: SmolLM3 OLMo-2-0425-1B OLMo-2-1124-7B Splits Each split (smollm3_3b, olmo2_1b, olmo2_7b) contains rows for model checkpoints and columns for benchmark scores (e.g., ASSIN2 RTE, ENEM… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-eval-logs-olmo2-smollm3.imagen<1K0 likes125 downloads7mo agoHugging Face20verify-ppt /smollm3-stack-v2-TypeScript0 likes121 downloads6mo agoHugging Face21lighteval /RULER-131072-SmolLM3-11T-32k-v1-remote-codetabular1K<n<10K0 likes103 downloads1y agoHugging Face22verify-ppt /smollm3-stack-v2-Java0 likes103 downloads6mo agoHugging Face23reasoning-cues /rollouts-smollm3-csail rollouts-smollm3-csail Model: HuggingFaceTB/SmolLM3-3B-Base and its intermediate checkpoints (smollm3_ladder/). Tokenizer: HuggingFaceTB/SmolLM3-3B-Base. Protocol: entropy screens: none + the model's top-20 beam nominees, first 30 MATH-train problems x 16 rollouts, budget 16,384, T 0.6, top-p 0.95, seed 20260819; code screens: MBPP beam nominees on HumanEval 164 x 32, budget 31,744, execution-graded copies included; smollm3_confirm: MATH-500 x 4 at 31,744 (none, p2_bare, p2_okay… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-smollm3-csail.0 likes103 downloads12d agoHugging Face24lighteval /RULER-131072-SmolLM3-SFT0 likes102 downloads1y agoHugging Face25bestdive /details_bestdive__SmolLM3-3B-SFT-Free-Course Smol course SFT evaluation - Kay Zheng Actual full GSM8K test evaluation of bestdive/SmolLM3-3B-SFT-Free-Course, adapter revision 0484e028b494d605a267050a949c9266edadd16b, merged with pinned SmolLM3-3B-Base before evaluation. Full 1319 test examples, zero-shot, original extractive_match: 0.4086429112964367 (stderr 0.013540639733342422). Free Google Colab T4, no paid HF Jobs; cost 0. lighteval 0.11.0, vLLM 0.10.1.1, Transformers 4.57.1, Python 3.12. Dataset-address correction… See the full description on the dataset page: https://huggingface.co/datasets/bestdive/details_bestdive__SmolLM3-3B-SFT-Free-Course.textn<1K0 likes101 downloads15d agoHugging Face26lighteval /RULER-65536-SmolLM3-SFT0 likes99 downloads1y agoHugging Face27lighteval /RULER-262144-SmolLM3-11T-32k-v1-remote-codetabular1K<n<10K0 likes97 downloads1y agoHugging Face28verify-ppt /smollm3-stack-v2-C-Sharp0 likes95 downloads6mo agoHugging Face29Harvard-DCML /tis-subset-datasets-SmolLM3-3B-Basetext100K<n<1M0 likes85 downloads8mo agoHugging Face30verify-ppt /smollm3-stack-v2-HTML0 likes82 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.