CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jordangong /the-stack-v2-smollm3 The Stack v2 — materialized source code Upstream dataset: bigcode/the-stack-v2 Exact upstream commit: e565caa3a78c2423bd374333a472b049eb090e47 Primary source-content endpoint: https://softwareheritage.s3.amazonaws.com/content/{blob_id} Configurations TypeScript Swift Ruby Rust Go Shell Jupyter_Notebook HTML Python Java JavaScript C C++ C-Sharp PHP SQL Markdown Added columns content: decoded source content download_error: null on successful… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/the-stack-v2-smollm3.texttext-generation1B<n<10B1 likes16k downloads15d agoHugging Face02Avelina /smollm-corpus-cleaned SmolLM-Corpus: Now shuffled and sharded (and Cleaned)! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.texttext-generation100M<n<1B2 likes11k downloads2y agoHugging Face03jordangong /jupyter-scripts-smollm3 The Stack v2 Jupyter Notebooks as Scripts This dataset contains script representations of the Jupyter notebooks in The Stack v2. It was created from the materialized Jupyter_Notebook split in jordangong/the-stack-v2-smollm3. The output schema follows the Jupyter-script schema used by bigcode/starcoderdata, but this release is not deduplicated, PII-filtered, or otherwise equivalent to StarCoderData's filtered split. Relationship to the SmolLM3 training mix This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.tabulartext-generation1M<n<10M0 likes1.1k downloads16d agoHugging Face04argilla-warehouse /apigen-smollm-trl-FC Dataset card for argilla-warehouse/apigen-smollm-trl-FC This dataset is a merge of argilla/Synth-APIGen-v0.1 and Salesforce/xlam-function-calling-60k, and was prepared for training using the script prepare_for_sft.py that can be found in the repository files. References @article{liu2024apigen, title={APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets}, author={Liu, Zuxin and Hoang, Thai and Zhang, Jianguo and Zhu, Ming and… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/apigen-smollm-trl-FC.texttext-generation100K<n<1M2 likes910 downloads2y agoHugging Face05enPurified /smollm-corpus-fineweb-edu-enPurified-openai-messages 📖 smollm-corpus-fineweb-edu-enPurified-openai-messages smollm-corpus-fineweb-edu-enPurified is a highly curated, "prose-first" subset of the fineweb-edu-dedup subset found in HuggingFaceTB/smollm-corpus. The enPurified collection is built on a specific philosophy: Specialization. While the original dataset is excellent for general pre-training, high-quality fluent English prose often gets diluted when mixed with syntax-heavy code, rigid math formulas, or low-information web junk.… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-fineweb-edu-enPurified-openai-messages.texttext-generation10M<n<100M1 likes343 downloads8mo agoHugging Face06enPurified /smollm-corpus-cosmopedia-v2-enPurified-openai-messages enPurified Collection: Smollm Corpus Cosmopedia V2] Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows. Purpose of the enPurified Collection The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.texttext-generation1M<n<10M1 likes288 downloads8mo agoHugging Face07david-thrower /smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming A corpus of high quality fine tuning data meant for fine tuning various HelixLM models Dataset Composition: A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ... Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning. Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.texttext-generation1M<n<10M0 likes171 downloads4mo agoHugging Face08zcamz /ai-vs-human-HuggingFaceTB-SmolLM2-1.7B-Instruct AI vs Human dataset on the CNN Daily mails Dataset Description This dataset showcases pairs of truncated articles and their respective completions, crafted either by humans or an AI language model. Each article was randomly truncated between 25% and 50% of its length. The language model was then tasked with generating a completion that mirrored the characters count of the original human-written continuation. Data Fields 'human': The original human-authored… See the full description on the dataset page: https://huggingface.co/datasets/zcamz/ai-vs-human-HuggingFaceTB-SmolLM2-1.7B-Instruct.texttext-classification1K<n<10K1 likes67 downloads2y agoHugging Face09ecreeth /1b-smollm-corpus SmolLM-Corpus — 1B Token Subset A curated 1-billion-token English pretraining corpus sampled from HuggingFaceTB/smollm-corpus, designed for training small language models (~20M parameters). Dataset Composition Source Ratio Tokens Documents FineWeb-Edu (dedup) 87% ~870M 849,577 Cosmopedia v2 13% ~130M 161,889 Total 100% ~1B 1,011,466 Rationale for the Split The 87/13 ratio mirrors the natural token distribution of the full… See the full description on the dataset page: https://huggingface.co/datasets/ecreeth/1b-smollm-corpus.texttext-generation1M<n<10M1 likes65 downloads4mo agoHugging Face10DSTI /traffic-accidents-reports-kd-smollm2-360M-7k Accident Reporting KD Dataset (One-Paragraph) Short description.A training/evaluation dataset for generating one-paragraph accident/incident reports from structured facts.This dataset mixes gold human targets from zBotta/traffic-accidents-reports-5k with teacher-generated soft targets produced by the model zBotta/smollm2-accident-reporter-360m to support knowledge distillation (KD) of a smaller student. Output style: a single paragraph, neutral tone, covering What, When, Where, Who… See the full description on the dataset page: https://huggingface.co/datasets/DSTI/traffic-accidents-reports-kd-smollm2-360M-7k.texttext-generation1K<n<10K1 likes31 downloads1y agoHugging Face11KBaba7 /SmoLLM-Dataset Dataset Card for SmoLLM-Dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/KBaba7/SmoLLM-Dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/KBaba7/SmoLLM-Dataset.texttext-generationn<1K0 likes30 downloads2y agoHugging Face12FatimaAfzal01 /smollm3-3b-base-blind-spots SmolLM3-3B-Base Blind Spots Dataset This dataset contains 10 test cases where I explored the failure modes of SmolLM3-3B-Base, a 3 billion parameter base language model released by HuggingFace in 2025. The goal was to find diverse cases where the model makes clearly incorrect or unexpected completions its "blind spots." Model Tested Model: HuggingFaceTB/SmolLM3-3B-Base Parameters: 3B Type: Base pretrained model License: Apache 2.0 How I Loaded the Model I… See the full description on the dataset page: https://huggingface.co/datasets/FatimaAfzal01/smollm3-3b-base-blind-spots.texttext-generationn<1K0 likes23 downloads7mo agoHugging Face13barlowa /smollm2-135m-abstention-posttrain-data SmolLM2-135M abstention post-training task Synthetic abstention-vs-fabrication task used in barlowa124/llm-posttraining and the checkpoints at barlowa/smollm2-135m-abstention-posttrain. Files data/ — the four training/eval splits. sft (960 prompt+completion), dpo (400 chosen/rejected pairs), eval (160 held-out entities), rl (400 GRPO prompts). Held-out entities never appear in train — contamination-free by construction. responses/ — raw model generations per… See the full description on the dataset page: https://huggingface.co/datasets/barlowa/smollm2-135m-abstention-posttrain-data.texttext-generation1K<n<10K0 likes21 downloads17h agoHugging Face14zcamz /ai-vs-human-HuggingFaceTB-SmolLM2-360M-Instruct AI vs Human dataset on the CNN Daily mails Dataset Description This dataset showcases pairs of truncated articles and their respective completions, crafted either by humans or an AI language model. Each article was randomly truncated between 25% and 50% of its length. The language model was then tasked with generating a completion that mirrored the characters count of the original human-written continuation. Data Fields 'human': The original human-authored… See the full description on the dataset page: https://huggingface.co/datasets/zcamz/ai-vs-human-HuggingFaceTB-SmolLM2-360M-Instruct.texttext-classification1K<n<10K1 likes20 downloads2y agoHugging Face15prometheus04 /matilda-smollm-mix-15b-gpt2 matilda-smollm-mix-15B-gpt2 15 B GPT-2-BPE tokens drawn from a 5:1 token-balanced mix of HuggingFaceTB/smollm-corpus: Source Share Tokens fineweb-edu-dedup 83.33 % 12.50 B cosmopedia-v2 16.67 % 2.50 B Total: 15,000,349,569 tokens across 151 shards (shard_*.bin, uint16, 100 M tokens per shard). The full SmolLM recipe is 75 / 15 / 10 fineweb-edu / cosmopedia-v2 / python-edu. python-edu was dropped because the HuggingFaceTB/smollm-corpus subset ships only blob_id… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/matilda-smollm-mix-15b-gpt2.tabulartext-generationn<1K1 likes16 downloads4mo agoHugging Face16Fu01978 /smollm_self_data smollm_self_data Dataset Description This dataset consists of 100 question-answer pairs generated entirely by the HuggingFaceTB/SmolLM2-135M-Instruct model. The dataset was created using a "self-prompting" approach, where the model was first asked to generate an interesting question, and then asked to provide an answer to that same question. The primary goal of this dataset is to serve as a base for fine-tuning of other small conversational models. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Fu01978/smollm_self_data.texttext-generationn<1K0 likes15 downloads7mo agoHugging Face17Alhibb /smollm2-blind-spots SmolLM2-1.7B Blind Spots Dataset This dataset contains 10 diverse examples where the SmolLM2-1.7B base model makes incorrect predictions or demonstrates "blind spots". Model Tested Model: SmolLM2-1.7B Parameters: 1.7 Billion Type: Base (Pre-trained) How to Load the Model The model was loaded using the transformers library in Python. import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "HuggingFaceTB/SmolLM2-1.7B" tokenizer =… See the full description on the dataset page: https://huggingface.co/datasets/Alhibb/smollm2-blind-spots.texttext-generationn<1K0 likes14 downloads7mo agoHugging Face18sapbot /gemma-3n-4b-distill-smollm2-360m-instruct-425xTrace of Gemma 3n 4B Distill SmolLM2 360M Instruct LLM by sapbot (me). Data count (Total: 425): English - 209 Russian - 216 Data is presented in ShareGPT format and each conversation split by newline. Note: This was added more as a "examples" of this model's outputs. Of course you will not distill a distilled model (I hope). Brought to you by sapbot from Romarchive texttext-generationn<1K0 likes14 downloads5mo agoHugging Face19NiazTahi /smollm3-base-blind-spots SmolLM3-3B-Base — Blind Spots Dataset Overview This dataset documents 12 diverse blind spots of the base language model HuggingFaceTB/SmolLM3-3B-Base (3 billion parameters, Apache-2.0, released 2025). Each row contains: Field Description id Integer index category Type of reasoning required input_prompt Partial text fed to the model expected_output Correct continuation model_output What SmolLM3-3B-Base actually generated notes Explanation of why the… See the full description on the dataset page: https://huggingface.co/datasets/NiazTahi/smollm3-base-blind-spots.texttext-generationn<1K0 likes13 downloads7mo agoHugging Face20aneeshadas02 /smollm3-3b-base-blind-spots SmolLM3-3B-Base Blind Spots Title & Overview A curated set of failure cases for HuggingFaceTB/SmolLM3-3B-Base, showcasing blind spots discovered while probing the 3B-parameter base pre-training checkpoint released in July 2025. Each entry captures a prompt, the expected aligned behaviour, and the model's actual output. The dataset illustrates common failure patterns observed when probing the base model without any instruction tuning, RLHF, or safety fine-tuning applied.… See the full description on the dataset page: https://huggingface.co/datasets/aneeshadas02/smollm3-3b-base-blind-spots.texttext-generationn<1K1 likes12 downloads7mo agoHugging Face21Dhruba461 /smollm3-3b-base-blindspots SmolLM3-3B-Base — Blind Spots Dataset This dataset contains 10 diverse input-output pairs where the base language model HuggingFaceTB/SmolLM3-3B-Base produces incorrect predictions under greedy decoding. Each row records the exact prompt fed to the model, the correct expected answer, and what the model actually generated — along with a description of the error type. Model Tested Field Value Model HuggingFaceTB/SmolLM3-3B-Base Parameters 3 billion… See the full description on the dataset page: https://huggingface.co/datasets/Dhruba461/smollm3-3b-base-blindspots.textquestion-answeringn<1K0 likes11 downloads7mo agoHugging Face22Shinzmann /smollm2-1.7b-blind-spots SmolLM2-1.7B Blind Spots Dataset A curated evaluation dataset documenting specific failure modes of HuggingFaceTB/SmolLM2-1.7B — a 1.7 billion parameter pretrained (base) language model. Each entry contains a completion-style prompt, the verified correct answer, and the model's actual incorrect output produced via deterministic greedy decoding. This dataset was created as part of the "Blind Spots of Frontier Models" technical challenge to systematically identify where small… See the full description on the dataset page: https://huggingface.co/datasets/Shinzmann/smollm2-1.7b-blind-spots.texttext-generationn<1K0 likes11 downloads7mo agoHugging Face23GoofyLM /smoll-cotThe dataset its still under development texttext-generationn<1K0 likes9 downloads1y agoHugging Face24Shah-4-8-1-2 /smollm2-1.7b-blindspots SmolLM2-1.7B Blind Spots Dataset A curated dataset of 12 diverse probe examples where the base language model HuggingFaceTB/SmolLM2-1.7B makes incorrect or unreliable predictions. Each row contains the raw prompt, the expected correct answer, the model's actual output (greedy decoding), the error category, and an explanation. Tested Model HuggingFaceTB/SmolLM2-1.7B Property Value Parameters 1.7 billion Type Pure base model (pretrained only — no… See the full description on the dataset page: https://huggingface.co/datasets/Shah-4-8-1-2/smollm2-1.7b-blindspots.texttext-generationn<1K0 likes9 downloads7mo agoHugging Face25mirackchuks /smollm2-blind-spots Model Tested HuggingFaceTB/SmolLM2-1.7B How I loaded it Used HuggingFace Transformers with AutoModelForCausalLM on Google Colab (T4 GPU, float16). Greedy decoding (do_sample=False) for reproducibility. View Colab Notebook Blind Spots Found The model struggled with: multi-step arithmetic, low-resource languages (Yoruba), African geographic knowledge, code generation, and logical reasoning. Fine-tuning Dataset Recommendation GSM8K / MATH for… See the full description on the dataset page: https://huggingface.co/datasets/mirackchuks/smollm2-blind-spots.texttext-generationn<1K0 likes8 downloads7mo agoHugging Face26habibahabchi /smollm3-base-blindspots SmolLM3-3B-Base Blind Spots Evaluation Dataset Dataset Summary This dataset documents 10 diverse failure cases discovered while evaluating HuggingFaceTB/SmolLM3-3B-Base, a 3-billion parameter decoder-only base language model released by Hugging Face in July 2025. The evaluation was conducted as part of the Fatima Fellowship technical challenge on Blind Spots of Frontier Models. Model Tested Model: HuggingFaceTB/SmolLM3-3B-Base Parameters: 3 billion… See the full description on the dataset page: https://huggingface.co/datasets/habibahabchi/smollm3-base-blindspots.tabulartext-generationn<1K0 likes8 downloads6mo agoHugging Face27Muhammad0981 /smollm2-blindspots Blind Spots Dataset for SmolLM2-1.7B Model Tested Model: SmolLM2-1.7B Parameters: 1.7B Release Date: February 2025 How I Loaded the Model from transformers import AutoModelForCausalLM, AutoTokenizer model_name = "HuggingFaceTB/SmolLM2-1.7B" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained( model_name, device_map="auto", torch_dtype="auto" ) def test_model(prompt, max_new_tokens=100): inputs =… See the full description on the dataset page: https://huggingface.co/datasets/Muhammad0981/smollm2-blindspots.texttext-generationn<1K1 likes8 downloads5mo agoHugging Face28hans1337 /smollm3-blindspots Blind Spots of SmolLM3-3B-Base This dataset documents systematic failure cases ("blind spots") observed when evaluating the SmolLM3-3B-Base model. The goal of this dataset is to identify patterns where a small base language model struggles with reasoning tasks that require precise symbolic or character-level manipulation. The dataset contains prompts where the model produces incorrect answers compared to the expected output. Model Tested Model:… See the full description on the dataset page: https://huggingface.co/datasets/hans1337/smollm3-blindspots.texttext-generationn<1K0 likes7 downloads7mo agoHugging Face29Sgobir /smollm3-blind-spots SmolLM3-3B-Base Blind Spots This dataset documents 10 failure cases observed while testing the model HuggingFaceTB/SmolLM3-3B-Base in Google Colab. Model Tested Model: HuggingFaceTB/SmolLM3-3B-Base How the Model Was Loaded The model was loaded in Google Colab using the transformers library with 4-bit quantization to run on limited GPU resources. Dataset Description This dataset contains 10 examples where the model produced incorrect outputs or failed… See the full description on the dataset page: https://huggingface.co/datasets/Sgobir/smollm3-blind-spots.texttext-generationn<1K0 likes7 downloads7mo agoHugging Face30BEE-spoke-data /smollm-corpus-pythongated smollm-corpus - python A version of the python-edu subset with the text added tabulartext-generation10M<n<100M0 likes6 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.