datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smol-smoltalk-Interaction-SFT
Dataset Card for ReactiveAI/Smol-Smoltalk Interaction SFT
Derived from HuggingFaceTB/smol-smoltalk. Made for Interaction Supervised Fine-Tuning of Reactive Transformer
Proof-of-Concept models, especially RxT-Beta.
Dataset Details
Dataset Description
Reactive Transformers are processing only the single interactions in real-time and using Short-Term Memory to store information from previous interactions.
Before the model is able to use it's memory, it has to be… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/smol-smoltalk-Interaction-SFT.smol-worldcup
🏟️ Smol AI WorldCup — SHIFT Benchmark
The world's first 5-axis evaluation framework for small language models.
Not just "how smart?" — but "how honest? how fast? how small? how efficient?"
🏟️ Leaderboard
huggingface.co/spaces/ginigen-ai/smol-worldcup
📊 Dataset
huggingface.co/datasets/ginigen-ai/smol-worldcup
🏅 ALL Bench
huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard
🏆 Official Ranking: WCS (WorldCup Score)
WCS = √( SHIFT × PIR_norm )… See the full description on the dataset page: https://huggingface.co/datasets/ginigen-ai/smol-worldcup.smol-koreantalkSmolLM2의 인스트럭션 훈련 데이터 HuggingFaceTB/smol-smoltalk를 한국어로 번역했어요.
smol-smoltalk-mini-Interaction-SFT
Dataset Card for ReactiveAI/Smol-Smoltalk-Mini Interaction SFT
Derived from HuggingFaceTB/smol-smoltalk (used 25% of train & test splits). Made for Interaction Supervised
Fine-Tuning of Reactive Transformer Proof-of-Concept models, especially RxT-Alpha-Mini (more info soon).
Full version available in ReactiveAI/smol-smoltalk-Interaction-SFT
Dataset Details
Dataset Description
Reactive Transformers are processing only the single interactions in… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/smol-smoltalk-mini-Interaction-SFT.smoke-openai-terra-batch-brasil-25-20260724-01
Smoke OpenAI Terra Batch — Brasil × 25 tasks
Run real de validação do fluxo matricial document_task_matrix, executada
sobre um único documento da Wikipédia em português com o título Brasil.
Cada uma das 25 tasks canônicas recebeu exatamente um slot inicial.
Resultado
status: completed
documentos: 1
pares planejados: 25
exemplos aceitos: 25
pares pulados: 0
pares esgotados: 0
resultados reais do backend: 27
retries com nova chamada: 2
backend: openai_api… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/smoke-openai-terra-batch-brasil-25-20260724-01.smoltalk2_everyday_conv_pt
SMOL Everyday Conversation PT
This dataset consists of a Portuguese version of the smoltalk_smollm3_everyday_conversations_no_think split of HuggingFaceTB/smoltalk2. The first two user turns were translated as well as the first assistant turn, the continuation of the conversation was generated using Gemma 3-27B.
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smoltalk2_everyday_conv_pt.smoothie-qwen3-8b-kr-self-driving-legal-dataset-v3
📚 Smoothie-Qwen3-8B-KR-Self-Driving-Legal Dataset v3 (DTRO Style)
대한민국 자율주행자동차법 파인튜닝을 위한 750건의 한국어 특화 데이터셋입니다.본 데이터셋은 기존 v1, v2 데이터셋 치명적인 "컨텍스트 소실(Context Forgetting)" 문제를 해결하기 위해 DTRO (Direct-To-Response Output) 스타일로 완전히 재구축되었습니다.
⚠️ 이전 데이터셋(v1, v2)의 문제점과 한계
기존 Alpaca 양식의 데이터셋은 모델 학습 시 다음과 같은 심각한 부작용을 낳았습니다.
1. 시스템 프롬프트 포이즈닝 (System Prompt Poisoning)
// 과거 데이터셋 (문제)
{
"instruction": "당신은 대한민국의 자율주행자동차법 전문가입니다. 정확한 법적 근거를...",
"input": "자율주행자동차의 정의는 무엇입니까? [출처: 관련… See the full description on the dataset page: https://huggingface.co/datasets/bluejude10/smoothie-qwen3-8b-kr-self-driving-legal-dataset-v3.smol-smoltalk-plus-reasoning-synthetic-data
Dataset Card for Smol-Smoltalk Plus Reasoning
This is a project to make a fork of HuggingFaceTB/smol-smoltalk which includes reasoning data generated using HuggingFaceTB/SmolLM2-1.7B-Instruct.
This is a work in progress. I ran a proof of concept on a small subset and will scale this up as I am able to.
Contributions to scale this up and complete this data are welcome, especially from those with access to more substantial GPU resources.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smol-smoltalk-plus-reasoning-synthetic-data.SmoLLM-Dataset
Dataset Card for SmoLLM-Dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/KBaba7/SmoLLM-Dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/KBaba7/SmoLLM-Dataset.smol-rewrite-PT
SMOL Rewrite PT
This dataset is the translated version of the smol-rewrite subset of the HuggingFaceTB/smoltalk.
This dataset includes an high-quality split used in the ramp down phase of the AMALIA's model post-training. The quality classification was done using google/gemma-4-31B-it.
Original Dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk
Note: This dataset comprises machine translated content and may contain translation errors or artifacts.
This… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smol-rewrite-PT.my_smolvla123456789smoltalk-no-refusals-augmented
smoltalk-no-refusals-augmented
A cleaned and augmented version of the smoltalk dataset, designed to minimize alignment priors and AI identity markers for research purposes.
Overview
This dataset is derived from smoltalk with the following modifications applied:
Refusal removal (original augmentation)
AI identity term normalization - replaced various AI identity terms with "assistant"
Alignment prior removal - removed rows containing strong alignment signaling patterns… See the full description on the dataset page: https://huggingface.co/datasets/EternalRecursion/smoltalk-no-refusals-augmented.Blind_Spots_Dataset_SmolLM3-3B-Base
SmolLM3-3B-Base Blind Spots Dataset
Dataset Summary
This dataset documents 13 failure cases of the HuggingFaceTB/SmolLM3-3B-Base model, a 3-billion parameter base language model pretrained on 11.2 trillion tokens. Each row contains a text completion prompt, the expected correct output, the model's actual output, and the category. The dataset spans 13 distinct categories in which the SmolLM3-3B-Base model fails to work as expected.
Model Tested
Model:… See the full description on the dataset page: https://huggingface.co/datasets/Yanmife/Blind_Spots_Dataset_SmolLM3-3B-Base.my-smol-ds
Dataset Card for my-smol-ds
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Binaryy/my-smol-ds/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Binaryy/my-smol-ds.smoltalk_azThis is part of the translated version of the original dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk
smollm3-3b-base-blindspots
SmolLM3-3B-Base — Blind Spots Dataset
This dataset contains 10 diverse input-output pairs where the base language model
HuggingFaceTB/SmolLM3-3B-Base
produces incorrect predictions under greedy decoding. Each row records the exact prompt fed
to the model, the correct expected answer, and what the model actually generated — along with
a description of the error type.
Model Tested
Field
Value
Model
HuggingFaceTB/SmolLM3-3B-Base
Parameters
3 billion… See the full description on the dataset page: https://huggingface.co/datasets/Dhruba461/smollm3-3b-base-blindspots.smol_reas_traces
Smol Reasoning Traces
This dataset contains extended Chain-of-Thought (CoT) reasoning traces generated during robustness evaluation experiments on Small Language Models (SLMs). It provides a detailed look at how models handle mathematical reasoning under various textual perturbations.
Performance Summary
1. GSM8K Standard Benchmark
Comparison of Pass@1 (Accuracy) and Pass@Any (at least one correct out of 16 traces).
Model
Samples
Pass@1 (Acc)
Pass@Any… See the full description on the dataset page: https://huggingface.co/datasets/saracandu/smol_reas_traces.evosynth-gsm8k-smoke-cerebras-gsm8k
EvoSynth — Genetic Algorithm Evolutionary
Synthetic math word problem dataset generated from gsm8k using evolutionary prompting. No model training was used — all samples are produced purely through inference-time evolutionary pressure.
Generation method
Genetic Algorithm Evolutionary: Genetic algorithm loop: initial population from the configured init model, then G generations of fitness-guided selection, crossover, and mutation. Best individuals accumulate across… See the full description on the dataset page: https://huggingface.co/datasets/pltops/evosynth-gsm8k-smoke-cerebras-gsm8k.smollm3-base-blindspots
SmolLM3-3B-Base Blind Spots Evaluation Dataset
Dataset Summary
This dataset documents 10 diverse failure cases discovered while evaluating
HuggingFaceTB/SmolLM3-3B-Base,
a 3-billion parameter decoder-only base language model released by Hugging Face in July 2025.
The evaluation was conducted as part of the Fatima Fellowship technical challenge on Blind Spots of Frontier Models.
Model Tested
Model: HuggingFaceTB/SmolLM3-3B-Base
Parameters: 3 billion… See the full description on the dataset page: https://huggingface.co/datasets/habibahabchi/smollm3-base-blindspots.smollm3-blindspots
Blind Spots of SmolLM3-3B-Base
This dataset documents systematic failure cases ("blind spots") observed
when evaluating the SmolLM3-3B-Base model.
The goal of this dataset is to identify patterns where a small base
language model struggles with reasoning tasks that require precise
symbolic or character-level manipulation.
The dataset contains prompts where the model produces incorrect answers
compared to the expected output.
Model Tested
Model:… See the full description on the dataset page: https://huggingface.co/datasets/hans1337/smollm3-blindspots.SmolDataEnvs
📈 SmolDataEnvs
5.5K+ RL tasks for hill-climbing small models in code and data science.
A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on.
Two runs over the same 5,000 tasks: shuffled against a curriculum ordered easiest to hardest.
Data-analysis tasks as a plain, load-and-go dataset: no runtime, no framework required. Each
row is one self-contained task: a real tabular dataset, a question about it, and a gold answer… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/SmolDataEnvs.
