CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ReactiveAI /smol-smoltalk-Interaction-SFT Dataset Card for ReactiveAI/Smol-Smoltalk Interaction SFT Derived from HuggingFaceTB/smol-smoltalk. Made for Interaction Supervised Fine-Tuning of Reactive Transformer Proof-of-Concept models, especially RxT-Beta. Dataset Details Dataset Description Reactive Transformers are processing only the single interactions in real-time and using Short-Term Memory to store information from previous interactions. Before the model is able to use it's memory, it has to be… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/smol-smoltalk-Interaction-SFT.texttext-generation1M<n<10M2 likes734 downloads1y agoHugging Face02ginigen-ai /smol-worldcup 🏟️ Smol AI WorldCup — SHIFT Benchmark The world's first 5-axis evaluation framework for small language models. Not just "how smart?" — but "how honest? how fast? how small? how efficient?" 🏟️ Leaderboard huggingface.co/spaces/ginigen-ai/smol-worldcup 📊 Dataset huggingface.co/datasets/ginigen-ai/smol-worldcup 🏅 ALL Bench huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard 🏆 Official Ranking: WCS (WorldCup Score) WCS = √( SHIFT × PIR_norm )… See the full description on the dataset page: https://huggingface.co/datasets/ginigen-ai/smol-worldcup.tabulartext-generationn<1K47 likes420 downloads7mo agoHugging Face03lemon-mint /smol-koreantalkSmolLM2의 인스트럭션 훈련 데이터 HuggingFaceTB/smol-smoltalk를 한국어로 번역했어요. textquestion-answering100K<n<1M14 likes151 downloads2y agoHugging Face04ReactiveAI /smol-smoltalk-mini-Interaction-SFT Dataset Card for ReactiveAI/Smol-Smoltalk-Mini Interaction SFT Derived from HuggingFaceTB/smol-smoltalk (used 25% of train & test splits). Made for Interaction Supervised Fine-Tuning of Reactive Transformer Proof-of-Concept models, especially RxT-Alpha-Mini (more info soon). Full version available in ReactiveAI/smol-smoltalk-Interaction-SFT Dataset Details Dataset Description Reactive Transformers are processing only the single interactions in… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/smol-smoltalk-mini-Interaction-SFT.textquestion-answering100K<n<1M0 likes66 downloads1y agoHugging Face05costadev00 /smoke-openai-terra-batch-brasil-25-20260724-01 Smoke OpenAI Terra Batch — Brasil × 25 tasks Run real de validação do fluxo matricial document_task_matrix, executada sobre um único documento da Wikipédia em português com o título Brasil. Cada uma das 25 tasks canônicas recebeu exatamente um slot inicial. Resultado status: completed documentos: 1 pares planejados: 25 exemplos aceitos: 25 pares pulados: 0 pares esgotados: 0 resultados reais do backend: 27 retries com nova chamada: 2 backend: openai_api… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/smoke-openai-terra-batch-brasil-25-20260724-01.texttext-generationn<1K0 likes61 downloads2mo agoHugging Face06amalia-llm /smoltalk2_everyday_conv_pt SMOL Everyday Conversation PT This dataset consists of a Portuguese version of the smoltalk_smollm3_everyday_conversations_no_think split of HuggingFaceTB/smoltalk2. The first two user turns were translated as well as the first assistant turn, the continuation of the conversation was generated using Gemma 3-27B. Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smoltalk2_everyday_conv_pt.texttext-generation1K<n<10K1 likes51 downloads3mo agoHugging Face07bluejude10 /smoothie-qwen3-8b-kr-self-driving-legal-dataset-v3 📚 Smoothie-Qwen3-8B-KR-Self-Driving-Legal Dataset v3 (DTRO Style) 대한민국 자율주행자동차법 파인튜닝을 위한 750건의 한국어 특화 데이터셋입니다.본 데이터셋은 기존 v1, v2 데이터셋 치명적인 "컨텍스트 소실(Context Forgetting)" 문제를 해결하기 위해 DTRO (Direct-To-Response Output) 스타일로 완전히 재구축되었습니다. ⚠️ 이전 데이터셋(v1, v2)의 문제점과 한계 기존 Alpaca 양식의 데이터셋은 모델 학습 시 다음과 같은 심각한 부작용을 낳았습니다. 1. 시스템 프롬프트 포이즈닝 (System Prompt Poisoning) // 과거 데이터셋 (문제) { "instruction": "당신은 대한민국의 자율주행자동차법 전문가입니다. 정확한 법적 근거를...", "input": "자율주행자동차의 정의는 무엇입니까? [출처: 관련… See the full description on the dataset page: https://huggingface.co/datasets/bluejude10/smoothie-qwen3-8b-kr-self-driving-legal-dataset-v3.textquestion-answeringn<1K0 likes41 downloads7mo agoHugging Face08david-thrower /smol-smoltalk-plus-reasoning-synthetic-data Dataset Card for Smol-Smoltalk Plus Reasoning This is a project to make a fork of HuggingFaceTB/smol-smoltalk which includes reasoning data generated using HuggingFaceTB/SmolLM2-1.7B-Instruct. This is a work in progress. I ran a proof of concept on a small subset and will scale this up as I am able to. Contributions to scale this up and complete this data are welcome, especially from those with access to more substantial GPU resources. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smol-smoltalk-plus-reasoning-synthetic-data.textquestion-answering1K<n<10K5 likes40 downloads2y agoHugging Face09KBaba7 /SmoLLM-Dataset Dataset Card for SmoLLM-Dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/KBaba7/SmoLLM-Dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/KBaba7/SmoLLM-Dataset.texttext-generationn<1K0 likes30 downloads2y agoHugging Face10amalia-llm /smol-rewrite-PT SMOL Rewrite PT This dataset is the translated version of the smol-rewrite subset of the HuggingFaceTB/smoltalk. This dataset includes an high-quality split used in the ramp down phase of the AMALIA's model post-training. The quality classification was done using google/gemma-4-31B-it. Original Dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk Note: This dataset comprises machine translated content and may contain translation errors or artifacts. This… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smol-rewrite-PT.textquestion-answering10K<n<100K0 likes30 downloads3mo agoHugging Face11Amanpatel81 /my_smolvla123456789text-classification1M<n<10M0 likes22 downloads1y agoHugging Face12EternalRecursion /smoltalk-no-refusals-augmented smoltalk-no-refusals-augmented A cleaned and augmented version of the smoltalk dataset, designed to minimize alignment priors and AI identity markers for research purposes. Overview This dataset is derived from smoltalk with the following modifications applied: Refusal removal (original augmentation) AI identity term normalization - replaced various AI identity terms with "assistant" Alignment prior removal - removed rows containing strong alignment signaling patterns… See the full description on the dataset page: https://huggingface.co/datasets/EternalRecursion/smoltalk-no-refusals-augmented.texttext-generation100K<n<1M0 likes21 downloads10mo agoHugging Face13Yanmife /Blind_Spots_Dataset_SmolLM3-3B-Base SmolLM3-3B-Base Blind Spots Dataset Dataset Summary This dataset documents 13 failure cases of the HuggingFaceTB/SmolLM3-3B-Base model, a 3-billion parameter base language model pretrained on 11.2 trillion tokens. Each row contains a text completion prompt, the expected correct output, the model's actual output, and the category. The dataset spans 13 distinct categories in which the SmolLM3-3B-Base model fails to work as expected. Model Tested Model:… See the full description on the dataset page: https://huggingface.co/datasets/Yanmife/Blind_Spots_Dataset_SmolLM3-3B-Base.question-answeringn<1K0 likes15 downloads7mo agoHugging Face14Binaryy /my-smol-ds Dataset Card for my-smol-ds This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/Binaryy/my-smol-ds/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Binaryy/my-smol-ds.texttext-generationn<1K0 likes14 downloads2y agoHugging Face15LocalDoc /smoltalk_azThis is part of the translated version of the original dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk tabulartranslation100K<n<1M1 likes13 downloads10mo agoHugging Face16Dhruba461 /smollm3-3b-base-blindspots SmolLM3-3B-Base — Blind Spots Dataset This dataset contains 10 diverse input-output pairs where the base language model HuggingFaceTB/SmolLM3-3B-Base produces incorrect predictions under greedy decoding. Each row records the exact prompt fed to the model, the correct expected answer, and what the model actually generated — along with a description of the error type. Model Tested Field Value Model HuggingFaceTB/SmolLM3-3B-Base Parameters 3 billion… See the full description on the dataset page: https://huggingface.co/datasets/Dhruba461/smollm3-3b-base-blindspots.textquestion-answeringn<1K0 likes13 downloads7mo agoHugging Face17saracandu /smol_reas_traces Smol Reasoning Traces This dataset contains extended Chain-of-Thought (CoT) reasoning traces generated during robustness evaluation experiments on Small Language Models (SLMs). It provides a detailed look at how models handle mathematical reasoning under various textual perturbations. Performance Summary 1. GSM8K Standard Benchmark Comparison of Pass@1 (Accuracy) and Pass@Any (at least one correct out of 16 traces). Model Samples Pass@1 (Acc) Pass@Any… See the full description on the dataset page: https://huggingface.co/datasets/saracandu/smol_reas_traces.texttext-generation10K<n<100K1 likes11 downloads9mo agoHugging Face18pltops /evosynth-gsm8k-smoke-cerebras-gsm8k EvoSynth — Genetic Algorithm Evolutionary Synthetic math word problem dataset generated from gsm8k using evolutionary prompting. No model training was used — all samples are produced purely through inference-time evolutionary pressure. Generation method Genetic Algorithm Evolutionary: Genetic algorithm loop: initial population from the configured init model, then G generations of fitness-guided selection, crossover, and mutation. Best individuals accumulate across… See the full description on the dataset page: https://huggingface.co/datasets/pltops/evosynth-gsm8k-smoke-cerebras-gsm8k.textquestion-answeringn<1K0 likes10 downloads5mo agoHugging Face19habibahabchi /smollm3-base-blindspots SmolLM3-3B-Base Blind Spots Evaluation Dataset Dataset Summary This dataset documents 10 diverse failure cases discovered while evaluating HuggingFaceTB/SmolLM3-3B-Base, a 3-billion parameter decoder-only base language model released by Hugging Face in July 2025. The evaluation was conducted as part of the Fatima Fellowship technical challenge on Blind Spots of Frontier Models. Model Tested Model: HuggingFaceTB/SmolLM3-3B-Base Parameters: 3 billion… See the full description on the dataset page: https://huggingface.co/datasets/habibahabchi/smollm3-base-blindspots.tabulartext-generationn<1K0 likes8 downloads6mo agoHugging Face20hans1337 /smollm3-blindspots Blind Spots of SmolLM3-3B-Base This dataset documents systematic failure cases ("blind spots") observed when evaluating the SmolLM3-3B-Base model. The goal of this dataset is to identify patterns where a small base language model struggles with reasoning tasks that require precise symbolic or character-level manipulation. The dataset contains prompts where the model produces incorrect answers compared to the expected output. Model Tested Model:… See the full description on the dataset page: https://huggingface.co/datasets/hans1337/smollm3-blindspots.texttext-generationn<1K0 likes8 downloads6mo agoHugging Face21FineEnvs /SmolDataEnvs 📈 SmolDataEnvs 5.5K+ RL tasks for hill-climbing small models in code and data science. A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on. Two runs over the same 5,000 tasks: shuffled against a curriculum ordered easiest to hardest. Data-analysis tasks as a plain, load-and-go dataset: no runtime, no framework required. Each row is one self-contained task: a real tabular dataset, a question about it, and a gold answer… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/SmolDataEnvs.tabularquestion-answering1K<n<10K1 likes14h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.