CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HPAI-BSC /MedQA-Mixtral-CoT Dataset Card for medqa-cot Synthetically enhanced responses to the medqa dataset using mixtral. Dataset Details Dataset Description To increase the quality of answers from the training splits of the MedQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the question… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MedQA-Mixtral-CoT.textmultiple-choice10K<n<100K9 likes802 downloads2y agoHugging Face02blue-blues /medical_cot Medical Question-Answering Dataset A comprehensive collection of medical questions and detailed answers, designed for training and evaluating medical question-answering systems. Dataset Description Overview This dataset contains medical questions with multiple-choice answers and detailed explanations. Each question presents a clinical scenario and requires medical knowledge to determine the correct diagnosis, treatment, or underlying mechanism. Data… See the full description on the dataset page: https://huggingface.co/datasets/blue-blues/medical_cot.textquestion-answering100K<n<1M6 likes598 downloads2y agoHugging Face03HPAI-BSC /headqa-cot-llama31 headqa-cot Synthetically enhanced responses to the HeadQA dataset. Used to train Aloe-Beta model. Dataset Details Dataset Description To increase the quality of answers from the training splits of the HeadQA dataset, we leverage Llama-3.1-70B-Instruct to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a hand-crafted list of… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/headqa-cot-llama31.textquestion-answering1K<n<10K2 likes407 downloads1y agoHugging Face04HPAI-BSC /MMLU-medical-cot-llama31 MMLU-medical-cot Synthetically enhanced responses to the medical-related questions of the auxiliary train set of the MMLU dataset. Used to train Aloe-Beta model. Dataset Details Dataset Description First, we use Llama-3.1-70B-Instruct to filter the medical-related questions of the auxiliary train set of the MMLU dataset. Next, we leverage Mixtral-8x7B to… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MMLU-medical-cot-llama31.textquestion-answering1K<n<10K6 likes393 downloads10mo agoHugging Face05amd /Cot-Drop LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/Cot-Drop.texttext-generation10K<n<100K1 likes369 downloads7mo agoHugging Face06HPAI-BSC /MedMCQA-Mixtral-CoT Dataset Card for medmcqa-cot Synthetically enhanced responses to the medmcqa dataset using mixtral. Dataset Details Dataset Description To increase the quality of answers from the training splits of the MedMCQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MedMCQA-Mixtral-CoT.textquestion-answering100K<n<1M4 likes271 downloads2y agoHugging Face07suitai /salabs-stem-deep-reasoning-cot-v13 🧪 SALabs Multi-Domain STEM Deep Reasoning & Chain-of-Thought (CoT) Corpus (v13.0) [!IMPORTANT] 💳 Click Here to Purchase Enterprise Commercial License ($2,500 USD) & Instant 31.7MB Master Archive DownloadInstant download of the full lossless master package containing all 1,816 JSONL reasoning records + 13 complete uncompressed text corpora (31.72 MB uncompressed total) + commercial license certificate. 🌟 Executive Summary The SALabs STEM Deep Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-stem-deep-reasoning-cot-v13.texttext-generation1K<n<10K1 likes213 downloads18d agoHugging Face08bugrabilge /Bilge-Turkish-CoT-50K Bilge: Turkish Chain-of-Thought Dataset (50K) 50,000 örneklik Türkçe Chain-of-Thought (CoT) reasoning fine-tuning veri seti. Bilge, Türkçe büyük dil modellerinin adım adım düşünme (reasoning) kapasitesini geliştirmek amacıyla hazırlanmış bir Chain-of-Thought veri setidir. Veri setindeki her örnek, modelin önce <think> blokları içinde görünür bir muhakeme süreci yürütmesini, ardından kullanıcıya yapılandırılmış ve detaylı bir cevap vermesini öğretmek üzere tasarlanmıştır. Bu… See the full description on the dataset page: https://huggingface.co/datasets/bugrabilge/Bilge-Turkish-CoT-50K.texttext-generation10K<n<100K9 likes185 downloads4mo agoHugging Face09HPAI-BSC /medqa-cot-llama31 medqa-cot-llama31 Synthetically enhanced responses to the MedQa dataset. Used to train Aloe-Beta model. Dataset Details Dataset Description To increase the quality of answers from the training splits of the MedQA dataset, we leverage Llama-3.1-70B-Instruct to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a hand-crafted… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/medqa-cot-llama31.textmultiple-choice10K<n<100K3 likes144 downloads10mo agoHugging Face10AlicanKiraz0 /Turkish-CoT-Instruct-Dataset 🇹🇷 Turkish CoT Instruct Dataset Türkçe Düşünme Zinciri (Chain-of-Thought) İçeren Talimat Veri Seti Bu veri seti, modellerin Türkçe adım adım akıl yürütme (reasoning) yeteneğini geliştirmek için hazırlanmıştır. Her örnekte model, cevabı vermeden önce <think> ... </think> etiketleri arasında tamamen Türkçe olarak adım adım düşünür, ardından ayrıntılı bir nihai cevap sunar (DeepSeek-R1 tarzı biçim). Örnek sayısı: 4.868 Dil: Türkçe Biçim: Sohbet (messages) — system / user /… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-CoT-Instruct-Dataset.texttext-generation1K<n<10K20 likes133 downloads2mo agoHugging Face11Jackrong /glm-4.7-multiturn-CoT glm-4.7-multiturn-CoT Dataset Summary glm-4.7-multiturn-CoT is a ShareGPT-style multi-turn reasoning distillation dataset generated with GLM-4.7 as the teacher model. This release focuses on preserving multi-turn dialogue continuity while injecting explicit chain-of-thought style responses in assistant turns. Key Features Multi-turn conversation format (human / gpt) Assistant responses stored as <think>...</think> + final answer Resume-safe distillation… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/glm-4.7-multiturn-CoT.texttext-generation1K<n<10K16 likes101 downloads7mo agoHugging Face12HAD653 /gsm8k-cot-120b 🚀 GSM8K-Teacher-CoT-120B (2025) High-Quality Short Chain-of-Thought Distillation Dataset Plain Text • No LaTeX • No ChatML • Deterministic Final Answers This dataset provides high-quality short Chain-of-Thought (CoT) reasoning generated by OpenAI gpt-oss-120b on the GSM8K benchmark. It is designed for small reasoning models (7B–14B). The dataset is: ✔ plain text ✔ concise and deterministic ✔ fully normalized ✔ free of LaTeX, ChatML, XML, Markdown ✔ optimized for tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/HAD653/gsm8k-cot-120b.textquestion-answering1K<n<10K4 likes100 downloads10mo agoHugging Face13HPAI-BSC /medmcqa-cot-llama31 medqa-cot-llama31 Synthetically enhanced responses to the MedMCQA dataset. Used to train Aloe-Beta model. Dataset Details Dataset Description To increase the quality of answers from the training splits of the MedMCQA dataset, we leverage Llama-3.1-70B-Instruct to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/medmcqa-cot-llama31.textmultiple-choice100K<n<1M2 likes99 downloads1y agoHugging Face14blythet /deepseek-v4-pro-math-cot-1k DeepSeek V4 Pro Math CoT 1K A small, high-signal supervised-fine-tuning (SFT) dataset of math reasoning traces. Problems were sampled from a Nemotron math problem set (originally sourced from StackExchange-Math and AoPS), answered by DeepSeek V4 Pro with thinking enabled at high reasoning effort, then independently reviewed by DeepSeek V4 Flash for correctness against the expected answer. Pathological reasoning traces (looping, run-away length, excessive Wait-style backtracking)… See the full description on the dataset page: https://huggingface.co/datasets/blythet/deepseek-v4-pro-math-cot-1k.tabulartext-generation1K<n<10K4 likes96 downloads5mo agoHugging Face15kaushik-harsh-99 /math-sft-solutions-no-cot Math SFT Solutions No CoT A cleaned mathematics supervised fine-tuning dataset containing: instruction → solution pairs mathematical proofs derivations olympiad-style solutions theorem reasoning stepwise mathematical explanations detailed final solutions This dataset was built specifically for mathematical supervised fine-tuning (SFT). Unlike many reasoning datasets, this release removes explicit chain-of-thought tags and hidden thinking traces while preserving high-quality… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/math-sft-solutions-no-cot.texttext-generation100K<n<1M5 likes94 downloads4mo agoHugging Face16alexfromapex /elementary_math_cot Elementary Math QA with Chain-of-Thought Dataset Description This dataset contains synthetically generated elementary math questions (addition, subtraction, multiplication, division, order of operations/PEMDAS, percentages, exponents, square roots, and averages), each paired with a step-by-step chain-of-thought (CoT) explanation and a final answer. Problems and their ground-truth answers are generated deterministically in Python, so every final answer is exact… See the full description on the dataset page: https://huggingface.co/datasets/alexfromapex/elementary_math_cot.textquestion-answering1K<n<10K0 likes93 downloads19d agoHugging Face17HPAI-BSC /PubmedQA-Mixtral-CoT Dataset Card for pubmedqa-cot Synthetically enhanced responses to the pubmedqa dataset using mixtral. Dataset Details Dataset Description To increase the quality of answers from the training splits of the PubMedQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/PubmedQA-Mixtral-CoT.textmultiple-choice100K<n<1M2 likes92 downloads2y agoHugging Face18kaushik-harsh-99 /math-sft-solutions-no-cot-v3 Math SFT Solutions No CoT V3 Math SFT Solutions No CoT V3 is a large-scale mathematics supervised fine-tuning (SFT) dataset designed for instruction tuning and mathematical capability adaptation. Version 3 substantially expands mathematical coverage while improving dataset quality through stronger filtering, cleaning, and supervision refinement. Unlike reasoning-heavy datasets, this release focuses on clean instruction → response pairs without hidden chain-of-thought style… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/math-sft-solutions-no-cot-v3.texttext-generation1M<n<10M5 likes78 downloads4mo agoHugging Face19andyc03 /PRISM-CoT PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality This repository contains the PRISM-CoT and PRISM-DPO datasets, which are key components of the PRISM (Principled Reasoning for Integrated Safety in Multimodality) framework. PRISM is a system2-like framework designed to align Vision-Language Models (VLMs) by embedding a structured, safety-aware reasoning process. PRISM-CoT is a dataset that teaches safety-aware chain-of-thought reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/andyc03/PRISM-CoT.textquestion-answering10K<n<100K0 likes70 downloads1y agoHugging Face20batuhanozkose /Rehber-CoT-Science 🧬 Rehber-CoT-Science: Turkish Scientific Reasoning Dataset Turkish Scientific Computational Reasoning (Chain-of-Thought) Dataset Multi-step scientific problem-solving dataset with verifiable Python code and detailed explanations Dataset • Author 📌 Changelog Eski sürümlere erişim: Branch menüsünden v1 seçebilirsiniz. Version Date Changes v2.0 24.12.2025 ✨ Yeni explained_answer alanı eklendi, Statistics domain eklendi, 712 örneğe genişletildi… See the full description on the dataset page: https://huggingface.co/datasets/batuhanozkose/Rehber-CoT-Science.textquestion-answering1K<n<10K4 likes64 downloads9mo agoHugging Face21Svngoku /AfriHist-CoT AfriHist-CoT Overview AfriHist-CoT is a dataset of question-answer pairs derived from African history books, created using a Chain-of-Thought (CoT) reasoning approach with the Gemini language model via OpenRouter. The dataset supports training and evaluating question-answering models, with a focus on African history and CoT reasoning. It is available in English and French, catering to both monolingual and multilingual applications. Dataset Description The… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/AfriHist-CoT.texttext-generation10K<n<100K3 likes63 downloads1y agoHugging Face22alexfromapex /simplemath-cot 🧮 SimpleMath-100k CoT A chain-of-thought (CoT) extension of the ProCreations/SimpleMath dataset. Every one of the 100 000 algebra / arithmetic problems is paired with a short, numbered reasoning trace (Step 1: … Step 2: …) that walks a language model from the problem statement to the known-correct answer. The traces in the Jupyter notebook are generated by Qwen3.8-27B and then post-processed to strip formatting noise, enforce sequential step numbering, and cap output at 1 000… See the full description on the dataset page: https://huggingface.co/datasets/alexfromapex/simplemath-cot.texttext-generationn<1K0 likes63 downloads19d agoHugging Face23Jianyuan1 /cot-dataDataset used in the paper Dyve: Thinking Fast and Slow for Dynamic Process Verification. Github: https://github.com/Jianyuan1/Dyve textquestion-answering100K<n<1M1 likes58 downloads2y agoHugging Face24hesamation /pubmedqa_cot_foltextquestion-answeringn<1K1 likes54 downloads1y agoHugging Face25waylake /ko-verified-cot ko-verified-cot 2,016 short Korean reasoning traces whose final answer was verified against a ground-truth answer key. Wrong reasoning was thrown away, not kept. How it was built Sample multiple-choice questions from the train split of KMMLU (45 subjects) — test split is never touched, so downstream evaluation stays clean. Ask the teacher (ox-alpha-free) for a short Korean chain of thought (3–4 sentences) ending in 정답: X. The teacher never sees the answer key.… See the full description on the dataset page: https://huggingface.co/datasets/waylake/ko-verified-cot.texttext-generation1K<n<10K0 likes53 downloads29d agoHugging Face26ceselder /concept-cot-conv-qa-full Concept CoT Conversational QA (Full) Conversational QA pairs about chain-of-thought reasoning traces, generated using DeepSeek v3.2 via OpenRouter. Designed for training activation oracles to answer natural language questions about what a model is doing during reasoning. Overview Total pairs: 10,499 Unique source entries: 7,496 (from 8,132 concept corpus entries) Source corpus: ceselder/concept-cot-corpus-full (concept_corpus/corpus_full.jsonl) Generator model:… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/concept-cot-conv-qa-full.textquestion-answering10K<n<100K0 likes45 downloads7mo agoHugging Face27HenryShan /Gemini-MMLU-CoT Gemini-MMLU-CoT: An Advanced Mathematical Reasoning Dataset A synthetic dataset of 7,000 multiple-choice mathematics questions featuring detailed Chain-of-Thought (CoT) reasoning. The content was generated by Google's Gemini model, with questions inspired by the mathematical sections of the MMLU (Massive Multitask Language Understanding) benchmark. Overview This dataset is designed for training and evaluating AI models on complex mathematical reasoning. It covers a wide… See the full description on the dataset page: https://huggingface.co/datasets/HenryShan/Gemini-MMLU-CoT.textquestion-answering1K<n<10K2 likes44 downloads11mo agoHugging Face28MichaelAnthony /lemonseed-cot-math-clean lemonseed-cot-math-clean LemonSeed — cleaned chain-of-thought math word problems. Contents cot_math_clean.jsonl (3530 rows) Format JSON Lines (.jsonl), one example per line. Provenance Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella. textquestion-answering1K<n<10K0 likes43 downloads29d agoHugging Face29AiCloser /sharegpt_cot_dataset A data set inspired by the "Reflection" method, three-dimensional thinking and cot This is the ShareGPT format. The data set was generated using multiple llm synthesis. textquestion-answering1K<n<10K7 likes42 downloads2y agoHugging Face30flatlander1024 /NuminaMath-CoT-filteredDataset that contains problems that appears in both QwQ-LongCoT-130K-cleaned and NuminaMath-CoT. There are approximately 100k problems where the solution is in plain-CoT manner. textquestion-answering100K<n<1M1 likes42 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.