CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bigcode /the-stack-smolgated Dataset Description A small subset (~0.1%) of the-stack dataset, each programming language has 10,000 random samples from the original dataset. The dataset has 2.6GB of text (code). Languages The dataset contains 30 programming languages: "assembly", "batchfile", "c++", "c", "c-sharp", "cmake", "css", "dockerfile", "fortran", "go", "haskell", "html", "java", "javascript", "julia", "lua", "makefile", "markdown", "perl", "php", "powershell", "python", "ruby", "rust"… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol.tabulartext-generation100K<n<1M95 likes27k downloads3y agoHugging Face02Avelina /smollm-corpus-cleaned SmolLM-Corpus: Now shuffled and sharded (and Cleaned)! This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming! The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu-cleaned repo. Dataset Structure The dataset is split into 24 subdirectories, with the first 23… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus-cleaned.texttext-generation100M<n<1B2 likes11k downloads2y agoHugging Face03bigcode /the-stack-smol-xl Dataset Description A small subset of the-stack dataset, with 87 programming languages, each has 10,000 random samples from the original dataset. Languages The dataset contains 87 programming languages: 'ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c', 'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir', 'elm', 'emacs-lisp','erlang'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol-xl.tabulartext-generation100K<n<1M11 likes7.9k downloads4y agoHugging Face04ginigen-ai /smol-worldcup 🏟️ Smol AI WorldCup — SHIFT Benchmark The world's first 5-axis evaluation framework for small language models. Not just "how smart?" — but "how honest? how fast? how small? how efficient?" 🏟️ Leaderboard huggingface.co/spaces/ginigen-ai/smol-worldcup 📊 Dataset huggingface.co/datasets/ginigen-ai/smol-worldcup 🏅 ALL Bench huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard 🏆 Official Ranking: WCS (WorldCup Score) WCS = √( SHIFT × PIR_norm )… See the full description on the dataset page: https://huggingface.co/datasets/ginigen-ai/smol-worldcup.tabulartext-generationn<1K47 likes420 downloads7mo agoHugging Face05enPurified /smollm-corpus-fineweb-edu-enPurified-openai-messages 📖 smollm-corpus-fineweb-edu-enPurified-openai-messages smollm-corpus-fineweb-edu-enPurified is a highly curated, "prose-first" subset of the fineweb-edu-dedup subset found in HuggingFaceTB/smollm-corpus. The enPurified collection is built on a specific philosophy: Specialization. While the original dataset is excellent for general pre-training, high-quality fluent English prose often gets diluted when mixed with syntax-heavy code, rigid math formulas, or low-information web junk.… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-fineweb-edu-enPurified-openai-messages.texttext-generation10M<n<100M1 likes338 downloads8mo agoHugging Face06enPurified /smollm-corpus-cosmopedia-v2-enPurified-openai-messages enPurified Collection: Smollm Corpus Cosmopedia V2] Updated on January 15th to remove more math, code, and low quality English. The dataset has now been pruned from 39.1M rows down to ~9M rows. Purpose of the enPurified Collection The enPurified dataset collection is an initiative to curate strict, high-quality English prose datasets for language modeling. While the open-source community provides extensive resources for code, mathematics, and multilingual data, this… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smollm-corpus-cosmopedia-v2-enPurified-openai-messages.texttext-generation1M<n<10M1 likes230 downloads8mo agoHugging Face07ChinaunicomSoftware /smoltalk-chinese-QwQ-Distrill smoltalk-chinese-QwQ-Distrill [中文] [English] 📖Technical Report smoltalk-chinese-QwQ-Distrill is a Chinese fine-tuning dataset constructed with reference to the SmolTalk-Chinese dataset. It aims to provide high-quality synthetic reasoning data support for training large language models (LLMs). The dataset consists entirely of synthetic data, comprising over 700,000 entries. It is specifically designed to enhance the performance of Chinese LLMs across various tasks… See the full description on the dataset page: https://huggingface.co/datasets/ChinaunicomSoftware/smoltalk-chinese-QwQ-Distrill.tabulartext-generation100K<n<1M3 likes202 downloads2y agoHugging Face08enPurified /smoltalk-creative-writing-enPurified-openai-messages 📖 SmolTalk-Creative-Writing-enPurified-openai-messages SmolTalk-Creative-Writing-enPurified is a highly curated, "prose-first" subset of the original collinear-ai/smoltalk-creative-writing dataset. The enPurified collection is built on a specific philosophy: Specialization. While the ecosystem has plenty of datasets for coding (StackOverflow, StarCoder) and mathematics (GSM8K), high-quality, fluent English prose often gets diluted when mixed with syntax-heavy code or rigid math… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smoltalk-creative-writing-enPurified-openai-messages.texttext-generation10K<n<100K2 likes115 downloads9mo agoHugging Face09ychen /Generated-Empathetic-Dialogues-v0.1-Smol Generated Empathetic Conversations v0.1 - Smol This is dataset contains 10K rows of multi-round empathetic conversations convering a diverse set of topics. Highlights Multi-round conversation It's not single-turn. The user and the assistant works together to gradually unfold the conversation. The average number of turns is 5, with a standard deviation of approximately 1.59 turns. A turn consists of two messages with one by the user, and another by the… See the full description on the dataset page: https://huggingface.co/datasets/ychen/Generated-Empathetic-Dialogues-v0.1-Smol.texttext-generation10K<n<100K4 likes66 downloads2y agoHugging Face10costadev00 /smoke-openai-terra-batch-brasil-25-20260724-01 Smoke OpenAI Terra Batch — Brasil × 25 tasks Run real de validação do fluxo matricial document_task_matrix, executada sobre um único documento da Wikipédia em português com o título Brasil. Cada uma das 25 tasks canônicas recebeu exatamente um slot inicial. Resultado status: completed documentos: 1 pares planejados: 25 exemplos aceitos: 25 pares pulados: 0 pares esgotados: 0 resultados reais do backend: 27 retries com nova chamada: 2 backend: openai_api… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/smoke-openai-terra-batch-brasil-25-20260724-01.texttext-generationn<1K0 likes61 downloads2mo agoHugging Face11sbussiso /SmolThinker-Synthetic-Reasoning SmolThinker Synthetic Reasoning Dataset 7,796 examples that teach a small instruct model to emit its reasoning inside literal <think> ... </think> blocks, so chat UIs that render collapsible reasoning (Open WebUI, Ollama, LM Studio) pick them up. Model-agnostic: nothing in the data names a specific model. Reasoning length scales with task difficulty, from one line on a greeting to 2,000+ characters on a hard question. from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/SmolThinker-Synthetic-Reasoning.texttext-generation10K<n<100K1 likes55 downloads2mo agoHugging Face12amalia-llm /smoltalk2_everyday_conv_pt SMOL Everyday Conversation PT This dataset consists of a Portuguese version of the smoltalk_smollm3_everyday_conversations_no_think split of HuggingFaceTB/smoltalk2. The first two user turns were translated as well as the first assistant turn, the continuation of the conversation was generated using Gemma 3-27B. Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smoltalk2_everyday_conv_pt.texttext-generation1K<n<10K1 likes51 downloads3mo agoHugging Face13bluejude10 /smoothie-qwen3-8b-kr-self-driving-legal-dataset-v5-cot 🇰🇷 자율주행법령 CoT 파인튜닝 데이터셋 v5 왜 이 데이터셋을 새로 만들었는가? 기존 dataset-v3 는 단순 질문-답변(Direct To Response, DTRO Style) 포맷으로 구성되어 있었습니다. // v3 포맷 (기존) { "instruction": "자율주행자동차란 무엇인가요?", "output": "자율주행자동차란 ..." } 이 방식으로 파인튜닝한 모델(v3)을 RAG 파이프라인과 결합하여 평가한 결과, 정답률 43% 로 순정 모델(90%)에 크게 뒤처지는 것이 확인되었습니다. 실패의 핵심 원인은 다음과 같습니다: 실패 원인 설명 템플릿 과적합 모델이 논리가 아닌 답변 패턴(아닙니다 + 설명)을 암기 RAG 컨텍스트 무시 학습된 내부 패턴이 외부 검색 문서를 압도 <think> 태그 미사용 Qwen3의 추론(Chain-of-Thought) 능력이 전혀 활성화되지 않음… See the full description on the dataset page: https://huggingface.co/datasets/bluejude10/smoothie-qwen3-8b-kr-self-driving-legal-dataset-v5-cot.texttext-generationn<1K0 likes44 downloads7mo agoHugging Face14Smoked-Salmon-s /empathetic_dialogues_ko Dataset Card for "한국어 일상 속 공감형 대화 데이터셋(멀티-턴)" Dataset Summary boostCamp AI Tech 5기 과정 중 NLP 12조 훈제연어들 팀의 최종 프로젝트에서 제작한 데이터입니다. 일상 속 다양한 상황에서 사용자와 챗봇 간의 대화를 담은 데이터셋 입니다. GPT4, GPT3.5-turbo로 제작된 합성데이터이며 싱글-턴, 2-턴, 3-턴 대화로 구성되어 있습니다. 답변은 [공감적 표현 - 일반적인 대화 - 관련된 질문] 의 형태를 가집니다. Generation Prompt Example(GPT3.5-turbo) Take a close look at the following example and Conditions. Create nine sessions that each of the session is ongoing conversation about a single… See the full description on the dataset page: https://huggingface.co/datasets/Smoked-Salmon-s/empathetic_dialogues_ko.texttext-generation10K<n<100K8 likes35 downloads3y agoHugging Face15genlab-1c /prism-smopPRISM-SMOP Датасет генераций кода 1С:Предприятие (BSL) с четырёхосевой разметкой качества по метрике SMOP 🤗 Hugging Face · Бенчмарк PRISM · Лицензия CC BY 4.0 · Метрика SMOP 1225 генераций кода 1С:Предприятие (BSL) от 35 нейросетей на 35 задачах бенчмарка PRISM, размеченных по четырём осям метрики SMOP — синтаксис, смысл, оптимальность, платформа — автоматическим оценщиком L1. Внутри — SFT-подвыборка из 222 решений открытых моделей, прошедших все скрытые тесты. 1225 BSL… See the full description on the dataset page: https://huggingface.co/datasets/genlab-1c/prism-smop.texttext-generation1K<n<10K1 likes32 downloads2mo agoHugging Face16amalia-llm /smol-rewrite-PT SMOL Rewrite PT This dataset is the translated version of the smol-rewrite subset of the HuggingFaceTB/smoltalk. This dataset includes an high-quality split used in the ramp down phase of the AMALIA's model post-training. The quality classification was done using google/gemma-4-31B-it. Original Dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk Note: This dataset comprises machine translated content and may contain translation errors or artifacts. This… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smol-rewrite-PT.textquestion-answering10K<n<100K0 likes30 downloads3mo agoHugging Face17amalia-llm /smol_summarize_pt SMOL Summarize PT This dataset is the translated version of the smol-summarize subset of the HuggingFaceTB/smoltalk. Original Dataset: https://huggingface.co/datasets/HuggingFaceTB/smoltalk Note: This dataset comprises machine translated content and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in the data mix used to post-train the AMALIA model. Citation If you use this dataset or… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/smol_summarize_pt.texttext-generation10K<n<100K0 likes26 downloads3mo agoHugging Face18FatimaAfzal01 /smollm3-3b-base-blind-spots SmolLM3-3B-Base Blind Spots Dataset This dataset contains 10 test cases where I explored the failure modes of SmolLM3-3B-Base, a 3 billion parameter base language model released by HuggingFace in 2025. The goal was to find diverse cases where the model makes clearly incorrect or unexpected completions its "blind spots." Model Tested Model: HuggingFaceTB/SmolLM3-3B-Base Parameters: 3B Type: Base pretrained model License: Apache 2.0 How I Loaded the Model I… See the full description on the dataset page: https://huggingface.co/datasets/FatimaAfzal01/smollm3-3b-base-blind-spots.texttext-generationn<1K0 likes25 downloads7mo agoHugging Face19codex-master /numina_smoltalk_mixturetexttext-generation100K<n<1M0 likes23 downloads2mo agoHugging Face20sapbot /gemma-3n-4b-distill-smollm2-360m-instruct-425xTrace of Gemma 3n 4B Distill SmolLM2 360M Instruct LLM by sapbot (me). Data count (Total: 425): English - 209 Russian - 216 Data is presented in ShareGPT format and each conversation split by newline. Note: This was added more as a "examples" of this model's outputs. Of course you will not distill a distilled model (I hope). Brought to you by sapbot from Romarchive texttext-generationn<1K0 likes19 downloads5mo agoHugging Face21prometheus04 /matilda-smollm-mix-15b-gpt2 matilda-smollm-mix-15B-gpt2 15 B GPT-2-BPE tokens drawn from a 5:1 token-balanced mix of HuggingFaceTB/smollm-corpus: Source Share Tokens fineweb-edu-dedup 83.33 % 12.50 B cosmopedia-v2 16.67 % 2.50 B Total: 15,000,349,569 tokens across 151 shards (shard_*.bin, uint16, 100 M tokens per shard). The full SmolLM recipe is 75 / 15 / 10 fineweb-edu / cosmopedia-v2 / python-edu. python-edu was dropped because the HuggingFaceTB/smollm-corpus subset ships only blob_id… See the full description on the dataset page: https://huggingface.co/datasets/prometheus04/matilda-smollm-mix-15b-gpt2.tabulartext-generationn<1K1 likes19 downloads4mo agoHugging Face22KaiLiu1998 /smooth_reading Smooth Reading This dataset accompanies the paper Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Understanding. license: mit texttext-generation100K<n<1M0 likes16 downloads5mo agoHugging Face23CooperBench /team-coop-smoke CooperBench Team → Coop (Qwen3.5-9B smoke) Placeholder / smoke dataset (1 trajectory). A single 2-agent cooperbench team run (lead + member, no protocol) reshaped into the 2-agent coop layout defined in cooperbench/CooperData PR #98. The full dataset is the canonical place where future team→coop conversions will land; this entry validates the converter and the publishing pipeline. Source Source run logs/qwen35-smoke-mini-team-noproto/ Source repo… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/team-coop-smoke.texttext-generationn<1K0 likes15 downloads4mo agoHugging Face24codex-master /balanced_smoltalk_basetexttext-generation100K<n<1M0 likes15 downloads2mo agoHugging Face25Pinkstack /syngen-reasoning-example-80-smoltalk1Reasoning generated using https://huggingface.co/Pinkstack/syngen-reasoning-0.6b (some of it was cutt off due to 8192 max length instead of 16k) based on the original smoltalk texttext-generationn<1K0 likes13 downloads11mo agoHugging Face26GoofyLM /smoll-cotThe dataset its still under development texttext-generationn<1K0 likes12 downloads1y agoHugging Face27aneeshadas02 /smollm3-3b-base-blind-spots SmolLM3-3B-Base Blind Spots Title & Overview A curated set of failure cases for HuggingFaceTB/SmolLM3-3B-Base, showcasing blind spots discovered while probing the 3B-parameter base pre-training checkpoint released in July 2025. Each entry captures a prompt, the expected aligned behaviour, and the model's actual output. The dataset illustrates common failure patterns observed when probing the base model without any instruction tuning, RLHF, or safety fine-tuning applied.… See the full description on the dataset page: https://huggingface.co/datasets/aneeshadas02/smollm3-3b-base-blind-spots.texttext-generationn<1K1 likes10 downloads7mo agoHugging Face28hans1337 /smollm3-blindspots Blind Spots of SmolLM3-3B-Base This dataset documents systematic failure cases ("blind spots") observed when evaluating the SmolLM3-3B-Base model. The goal of this dataset is to identify patterns where a small base language model struggles with reasoning tasks that require precise symbolic or character-level manipulation. The dataset contains prompts where the model produces incorrect answers compared to the expected output. Model Tested Model:… See the full description on the dataset page: https://huggingface.co/datasets/hans1337/smollm3-blindspots.texttext-generationn<1K0 likes8 downloads6mo agoHugging Face29Amin-AQ /smollm2-dpo-preferencesgated DPO Preferences Dataset (Restricted Access) Access Policy (Restricted) This dataset repo is public with manual gated access. Only approved users (from lums.edu.pk) will be granted access. Intended Use Preference optimization / DPO experiments for model alignment. Research and controlled evaluation. Out-of-Scope Use Any harmful, abusive, or policy-violating application. Safety-critical deployment without additional safeguards. Files… See the full description on the dataset page: https://huggingface.co/datasets/Amin-AQ/smollm2-dpo-preferences.texttext-generation1K<n<10K0 likes4 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.