CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lfaviate /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K3 likes598 downloads7mo agoHugging Face02suitai /salabs-stem-deep-reasoning-cot-v13 🧪 SALabs Multi-Domain STEM Deep Reasoning & Chain-of-Thought (CoT) Corpus (v13.0) [!IMPORTANT] 💳 Click Here to Purchase Enterprise Commercial License ($2,500 USD) & Instant 31.7MB Master Archive DownloadInstant download of the full lossless master package containing all 1,816 JSONL reasoning records + 13 complete uncompressed text corpora (31.72 MB uncompressed total) + commercial license certificate. 🌟 Executive Summary The SALabs STEM Deep Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-stem-deep-reasoning-cot-v13.texttext-generation1K<n<10K1 likes216 downloads19d agoHugging Face03Magpie-Align /Magpie-Reasoning-V2-250K-CoT-Llama3 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Reasoning-V2-250K-CoT-Llama3.tabulartext-generation100K<n<1M11 likes188 downloads2y agoHugging Face04ansulev /opus-4.7-reasoning-cot-4.8k Opus 4.7 Chain-of-Thought Reasoning 2,405 chain-of-thought reasoning traces produced by claude-opus-4-7 on hard reasoning prompts spanning math, science, and formal subjects. Each sample is a problem → <think> block → polished answer pair, where the <think> block contains Opus 4.7's full working (Restatement → Approach → Step-by-step derivation → Verification) and the post-</think> answer is written as a standalone lesson starting with the result in bold. How the… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/opus-4.7-reasoning-cot-4.8k.texttext-generation1K<n<10K5 likes102 downloads5mo agoHugging Face05eddieran /opus-4.7-reasoning-cot Opus 4.7 Chain-of-Thought Reasoning 2,405 chain-of-thought reasoning traces produced by claude-opus-4-7 on hard reasoning prompts spanning math, science, and formal subjects. Each sample is a problem → <think> block → polished answer pair, where the <think> block contains Opus 4.7's full working (Restatement → Approach → Step-by-step derivation → Verification) and the post-</think> answer is written as a standalone lesson starting with the result in bold. How the data… See the full description on the dataset page: https://huggingface.co/datasets/eddieran/opus-4.7-reasoning-cot.texttext-generation1K<n<10K15 likes74 downloads5mo agoHugging Face06miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes71 downloads1y agoHugging Face07expertdata-factory /cybersecurity-reasoning-cot-v1 🛡️ Expert Cybersecurity Reasoning Dataset (CoT) This dataset contains 89 high-fidelity, expert-verified reasoning records focusing on complex cybersecurity attack vectors. It is designed specifically for fine-tuning Large Language Models (LLMs) on sophisticated security analysis and threat logic. 💎 Key Highlights Niche Rarity 1.0: Covers rare and emerging threats with zero prior representation in open-source datasets. Advanced Vectors: Includes detailed reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/cybersecurity-reasoning-cot-v1.tabulartext-generationn<1K2 likes63 downloads7mo agoHugging Face08a13905873166 /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740… See the full description on the dataset page: https://huggingface.co/datasets/a13905873166/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K1 likes54 downloads9d agoHugging Face09abhinavdread /CoT-Scientific-RAG-Reasoning CoT-Scientific-RAG-Reasoning This dataset is designed for fine-tuning Large Language Models (specifically Qwen-series) to perform complex reasoning over scientific and technical documents using Chain-of-Thought (CoT). Dataset Description The dataset contains instructions and scientific contexts (Medical Imaging, Autonomous Driving, VLA Frameworks) where the model is required to generate a reasoning trace before providing the final answer. Format: JSONL Logic: All outputs… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/CoT-Scientific-RAG-Reasoning.texttext-generation1K<n<10K0 likes41 downloads9mo agoHugging Face10DuoNeural /cot-reasoning-2k DuoNeural CoT Reasoning Dataset (2K) A compact, high-quality chain-of-thought reasoning dataset generated for supervised fine-tuning (SFT). All 2,151 examples are quality-scored 5/5 and focus on explicit step-by-step reasoning traces. Benchmark Results Fine-tuned Qwen2.5-1.5B-Instruct on this dataset (3 epochs, LoRA rank 16, ~36 min on RTX 3090): Metric Baseline Post-SFT Δ Absolute Δ Relative GSM8K (flexible-extract) 0.3177 0.4890 +17.1pp +53.9% GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/cot-reasoning-2k.texttext-generation1K<n<10K1 likes39 downloads5mo agoHugging Face113amthoughts /Game_Reasoning_CoT 🎮 Game Reasoning CoT (Chain-of-Thought) Dataset Overview Game Reasoning CoT is a specialized dataset containing 551 records designed to fine-tune and evaluate LLMs on complex strategic decision-making and logical reasoning within gaming contexts. 📊 Dataset Statistics Total Samples: 551 Format: JSONL Categories: Chess, game_intelligence, Texas Hold'em, Blackjack, Roulette, Uno, Backgammon, Go Difficulty: {'hard': 522, 'medium': 29} 📊 Performance… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/Game_Reasoning_CoT.texttext-generationn<1K1 likes36 downloads4mo agoHugging Face12mkd-minju /keural-v2-cot-reasoning Reasoning / Chain-of-Thought (Area 4) — Korean SFT Dataset Prep 상태: 비공개 스테이징 (private) — 제2자 감사 전, 공개 배포 대상 아님 출처 원본: nvidia/OpenMathReasoning (cot split) 커밋 해시: d3d08664755704f422af97d43a7ff0ded4bd95df 라이선스: CC-BY-4.0 (태그와 본문 일치, "License/Terms of Use: cc-by-4.0") 생성 모델: DeepSeek-R1(샘플 중 다수), QwQ-32B — 둘 다 오픈 웨이트 모델, 독점 모델 ToS 리스크 없음 언어: 영어 (지침서 §1.4 정책에 따라 번역 없이 영어 그대로 사용) 출처 구성 (problem_source) 문제(질문) 출처는 대부분 AoPS(Art of Problem Solving) 포럼… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-cot-reasoning.texttext-generation10K<n<100K0 likes30 downloads25d agoHugging Face13domofon /fake_news_cot_reasoning Fake News Chain-of-Thought Reasoning Dataset This dataset contains 9,500 news articles with Chain-of-Thought (CoT) reasoning explanations for why each article is classified as fake or real news. Dataset Description Each record contains natural language reasoning that explains the classification decision, generated using Qwen 2.5 1.5B model via llama.cpp inference. Features Field Type Description title string News article headline input string Full… See the full description on the dataset page: https://huggingface.co/datasets/domofon/fake_news_cot_reasoning.texttext-classification1K<n<10K0 likes26 downloads9mo agoHugging Face14korra141 /fingpt-dow30-cot-reasoning FinGPT Forecaster DOW30 — Chain-of-Thought Dataset Weekly stock price movement predictions for DOW-30 constituents, augmented with Chain-of-Thought (CoT) reasoning generated via GPT-4. Built on top of the base dataset: FinGPT/fingpt-forecaster-dow30-202305-202405 Schema Field Description symbol DOW-30 ticker (e.g. AXP, MSFT) period Forecast week, e.g. 2023-05-14 to 2023-05-21 prompt Full instruction prompt fed to the model answer CoT reasoning +… See the full description on the dataset page: https://huggingface.co/datasets/korra141/fingpt-dow30-cot-reasoning.texttext-generation1K<n<10K0 likes22 downloads3mo agoHugging Face15lucaswychan /CoT-Moderate-Reasoning-Embedding Do Reasoning Models Enhance Embedding Models? Introduction This is the dataset used to evaluate the model similarity with the Hierarchical Representation Similarity Analysis (HRSA) framework in the paper Do Reasoning Models Enhance Embedding Models?. To verify if the latent manifold is preserved even within reasoning trajectories, we construct a Chain-of-Thought (CoT) dataset. Unlike standard semantic datasets, this corpus… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/CoT-Moderate-Reasoning-Embedding.texttext-generationn<1K0 likes16 downloads7mo agoHugging Face16wop /Extreme-Reasoning-CoT Extreme Reasoning Still to be updated A dataset for Extra-Heavy Chain of Thought reasoning. AI assistants dont think good enough, this dataset is here to fix it. Elite level quality All rows are human supervised. textquestion-answeringn<1K2 likes16 downloads7mo agoHugging Face17mkd-minju /keural-v2-cot-reasoning-v2 Reasoning / Chain-of-Thought (Area 4, v2) — Korean SFT Dataset Prep 상태: 비공개 스테이징(private) — §3 처리(1~8번, 최종 인코딩 포함) 전부 완료. 제2자 감사 전, 공개 배포 대상 아님. 이 v2는 §3 처리를 새로 검증하며 진행한 최종 버전입니다(2026-08-10). v1(원본 problem/generated_solution 필드 그대로)과 달리, DeepSeek-V4-Flash-0731 학습용 최종 텍스트(text 필드)로 인코딩까지 완료됐습니다. 출처 원본: nvidia/OpenMathReasoning (cot split) 커밋 해시: d3d08664755704f422af97d43a7ff0ded4bd95df 라이선스: CC-BY-4.0 (태그와 본문 일치) 생성 모델: DeepSeek-R1(다수), QwQ-32B — 둘 다 오픈 웨이트 모델 언어:… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-cot-reasoning-v2.texttext-generation10K<n<100K0 likes16 downloads1mo agoHugging Face18mattwesney /CoT_Reasoning_Cookinggated Description: Embark on a flavorful journey into the intricate realm of culinary reasoning with the "CoT_Cooking_Reasoning" dataset. This open-source resource (MIT licensed) offers a carefully curated collection of question-and-answer pairs designed to train AI models in grasping the subtle yet significant nuances of culinary processes, ingredient relationships, and cooking time calculations. This dataset explores a wide range of culinary scenarios, from basic ingredient preparation and recipe… See the full description on the dataset page: https://huggingface.co/datasets/mattwesney/CoT_Reasoning_Cooking.texttext-generation1K<n<10K10 likes15 downloads1y agoHugging Face19lucaswychan /CoT-Hard-Reasoning-Embedding Do Reasoning Models Enhance Embedding Models? Introduction This is the dataset used to evaluate the model similarity with the Hierarchical Representation Similarity Analysis (HRSA) framework in the paper Do Reasoning Models Enhance Embedding Models?. To verify if the latent manifold is preserved even within reasoning trajectories, we construct a Chain-of-Thought (CoT) dataset. Unlike standard semantic datasets, this corpus… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/CoT-Hard-Reasoning-Embedding.texttext-generationn<1K0 likes14 downloads7mo agoHugging Face20lucaswychan /CoT-Easy-Reasoning-Embedding Do Reasoning Models Enhance Embedding Models? Introduction This is the dataset used to evaluate the model similarity with the Hierarchical Representation Similarity Analysis (HRSA) framework in the paper Do Reasoning Models Enhance Embedding Models?. To verify if the latent manifold is preserved even within reasoning trajectories, we construct a Chain-of-Thought (CoT) dataset. Unlike standard semantic datasets, this corpus… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/CoT-Easy-Reasoning-Embedding.texttext-generationn<1K0 likes10 downloads7mo agoHugging Face21bitwikiorg /schema_cot_reasoninggated ⊙ Prompt Programs for Agentic Reasoning Programmable task-dependent COTs for agentic reasoning. A 100-row seed dataset for programmable cognition. Each row defines: Prompt template + input binding + explanation + task-dependent reasoning program Pipeline: intake → binding → procedure → output Schema Column Meaning ID Stable row ID Name Task name Prompt Prompt template using {{VARIABLE}} Expression Input binding using $.path Explanation Binding… See the full description on the dataset page: https://huggingface.co/datasets/bitwikiorg/schema_cot_reasoning.texttext-generationn<1K1 likes5 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.