datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
salabs-stem-deep-reasoning-cot-v13
🧪 SALabs Multi-Domain STEM Deep Reasoning & Chain-of-Thought (CoT) Corpus (v13.0)
[!IMPORTANT]
💳 Click Here to Purchase Enterprise Commercial License ($2,500 USD) & Instant 31.7MB Master Archive DownloadInstant download of the full lossless master package containing all 1,816 JSONL reasoning records + 13 complete uncompressed text corpora (31.72 MB uncompressed total) + commercial license certificate.
🌟 Executive Summary
The SALabs STEM Deep Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-stem-deep-reasoning-cot-v13.SmartCode-Fable-5-Distill-CoT-Reasoning-1000x
🚀 SmartCode Fable 5 Distill CoT Reasoning 1000x
🕯️ One like or download equals a prayer for my bank account.
⚡ Overview
1,000 elite reasoning sequences from the brand new Fable 5. This is the ultimate bridge for injecting frontier intelligence into high-performance small models.
🔥 The Secret Sauce: Adaptive Reasoning
Most datasets fail by cramming massive, unusable complexity into small models. This dataset is different. We use Adaptive… See the full description on the dataset page: https://huggingface.co/datasets/mfielding92/SmartCode-Fable-5-Distill-CoT-Reasoning-1000x.CoT-Scientific-RAG-Reasoning
CoT-Scientific-RAG-Reasoning
This dataset is designed for fine-tuning Large Language Models (specifically Qwen-series) to perform complex reasoning over scientific and technical documents using Chain-of-Thought (CoT).
Dataset Description
The dataset contains instructions and scientific contexts (Medical Imaging, Autonomous Driving, VLA Frameworks) where the model is required to generate a reasoning trace before providing the final answer.
Format: JSONL
Logic: All outputs… See the full description on the dataset page: https://huggingface.co/datasets/abhinavdread/CoT-Scientific-RAG-Reasoning.cot-reasoning-2k
DuoNeural CoT Reasoning Dataset (2K)
A compact, high-quality chain-of-thought reasoning dataset generated for supervised fine-tuning (SFT). All 2,151 examples are quality-scored 5/5 and focus on explicit step-by-step reasoning traces.
Benchmark Results
Fine-tuned Qwen2.5-1.5B-Instruct on this dataset (3 epochs, LoRA rank 16, ~36 min on RTX 3090):
Metric
Baseline
Post-SFT
Δ Absolute
Δ Relative
GSM8K (flexible-extract)
0.3177
0.4890
+17.1pp
+53.9%
GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/cot-reasoning-2k.Game_Reasoning_CoT
🎮 Game Reasoning CoT (Chain-of-Thought) Dataset
Overview
Game Reasoning CoT is a specialized dataset containing 551 records designed to fine-tune and evaluate LLMs on complex strategic decision-making and logical reasoning within gaming contexts.
📊 Dataset Statistics
Total Samples: 551
Format: JSONL
Categories: Chess, game_intelligence, Texas Hold'em, Blackjack, Roulette, Uno, Backgammon, Go
Difficulty: {'hard': 522, 'medium': 29}
📊 Performance… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/Game_Reasoning_CoT.keural-v2-cot-reasoning
Reasoning / Chain-of-Thought (Area 4) — Korean SFT Dataset Prep
상태: 비공개 스테이징 (private) — 제2자 감사 전, 공개 배포 대상 아님
출처
원본: nvidia/OpenMathReasoning (cot split)
커밋 해시: d3d08664755704f422af97d43a7ff0ded4bd95df
라이선스: CC-BY-4.0 (태그와 본문 일치, "License/Terms of Use: cc-by-4.0")
생성 모델: DeepSeek-R1(샘플 중 다수), QwQ-32B — 둘 다 오픈 웨이트 모델, 독점 모델 ToS 리스크 없음
언어: 영어 (지침서 §1.4 정책에 따라 번역 없이 영어 그대로 사용)
출처 구성 (problem_source)
문제(질문) 출처는 대부분 AoPS(Art of Problem Solving) 포럼… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-cot-reasoning.EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-details
Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT
Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-details.CoT_Reasoning_The_Ancient_Past
Description:
Journey through the millennia and explore the world of ancient civilizations with the "CoT_Reasoning_The_Ancient_Past" dataset. This open-source resource (MIT licensed) offers a carefully curated collection of question-and-answer pairs designed to train AI models in grasping the subtle yet significant nuances of historical causation, societal structures, cultural developments, and the logical steps involved in interpreting evidence from the ancient past. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/mattwesney/CoT_Reasoning_The_Ancient_Past.spider-cot-reasoningSpider-Reasoning dataset
spider-cot-reasoning is a dataset based on the Spider Text-to-SQL dataset (https://huggingface.co/datasets/xlangai/spider) (https://yale-lily.github.io/spider). This dataset augments the Spider dataset with reasoning steps generated by gpt-4o from OpenAI.
Specifically, the question as well as the gold SQL is given to gpt-4o, and gpt-4o is then prompted to generated reasoning steps that could have been taken to reach the gold SQL solution.
Dataset details
This data set… See the full description on the dataset page: https://huggingface.co/datasets/NyanDoggo/spider-cot-reasoning.bird-cot-reasoninglogic_reasoning_cot_zhCoT_Reasoning_Cooking
Description:
Embark on a flavorful journey into the intricate realm of culinary reasoning with the "CoT_Cooking_Reasoning" dataset. This open-source resource (MIT licensed) offers a carefully curated collection of question-and-answer pairs designed to train AI models in grasping the subtle yet significant nuances of culinary processes, ingredient relationships, and cooking time calculations. This dataset explores a wide range of culinary scenarios, from basic ingredient preparation and recipe… See the full description on the dataset page: https://huggingface.co/datasets/mattwesney/CoT_Reasoning_Cooking.kilo_cot_reasoningCoT-Moderate-Reasoning-Embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to evaluate the model similarity with the Hierarchical Representation Similarity Analysis (HRSA) framework in the paper Do Reasoning Models Enhance Embedding Models?. To verify if the latent manifold is preserved even within reasoning trajectories, we construct a Chain-of-Thought (CoT) dataset. Unlike standard semantic datasets, this corpus… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/CoT-Moderate-Reasoning-Embedding.Extreme-Reasoning-CoT
Extreme Reasoning
Still to be updated
A dataset for Extra-Heavy Chain of Thought reasoning.
AI assistants dont think good enough, this dataset is here to fix it.
Elite level quality
All rows are human supervised.
keural-v2-cot-reasoning-v2
Reasoning / Chain-of-Thought (Area 4, v2) — Korean SFT Dataset Prep
상태: 비공개 스테이징(private) — §3 처리(1~8번, 최종 인코딩 포함) 전부 완료. 제2자 감사 전, 공개 배포 대상 아님.
이 v2는 §3 처리를 새로 검증하며 진행한 최종 버전입니다(2026-08-10). v1(원본 problem/generated_solution 필드 그대로)과 달리, DeepSeek-V4-Flash-0731 학습용 최종 텍스트(text 필드)로 인코딩까지 완료됐습니다.
출처
원본: nvidia/OpenMathReasoning (cot split)
커밋 해시: d3d08664755704f422af97d43a7ff0ded4bd95df
라이선스: CC-BY-4.0 (태그와 본문 일치)
생성 모델: DeepSeek-R1(다수), QwQ-32B — 둘 다 오픈 웨이트 모델
언어:… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/keural-v2-cot-reasoning-v2.CoT-Hard-Reasoning-Embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to evaluate the model similarity with the Hierarchical Representation Similarity Analysis (HRSA) framework in the paper Do Reasoning Models Enhance Embedding Models?. To verify if the latent manifold is preserved even within reasoning trajectories, we construct a Chain-of-Thought (CoT) dataset. Unlike standard semantic datasets, this corpus… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/CoT-Hard-Reasoning-Embedding.cleand_moremilk_CoT_Reasoning_Quantom_Physics_And_Computing元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Quantom_Physics_And_Computing
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Quantom_Physics_And_Computing
データ件数: 2,862
平均トークン数: 1,110
最大トークン数: 2,334
合計トークン数: 3,175,666
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 15.5 MB
加工内容:
メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。
難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Quantom_Physics_And_Computing.CoT-Easy-Reasoning-Embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to evaluate the model similarity with the Hierarchical Representation Similarity Analysis (HRSA) framework in the paper Do Reasoning Models Enhance Embedding Models?. To verify if the latent manifold is preserved even within reasoning trajectories, we construct a Chain-of-Thought (CoT) dataset. Unlike standard semantic datasets, this corpus… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/CoT-Easy-Reasoning-Embedding.cot-oracle-reasoning-termination-balancedcleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Scientific_Discovery_and_Research
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Scientific_Discovery_and_Research
データ件数: 3,733
平均トークン数: 1,193
最大トークン数: 2,489
合計トークン数: 4,453,517
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 23.2 MB
加工内容:
メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。
難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research.EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-V2-ORPO-details
Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT-V2-ORPO
Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT-V2-ORPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-V2-ORPO-details.CoT_Temporal_Reasoning_Dataset
Description:
Embark on a journey into the intricate realm of temporal reasoning with the "CoT_Temporal_Reasoning" dataset. This open-source resource (MIT licensed) offers a carefully curated collection of question-and-answer pairs designed to train AI models in grasping the subtle yet significant nuances of temporal relationships, event sequencing, and duration calculations. This dataset explores a wide range of scenarios, from basic date arithmetic and event ordering to more complex problems… See the full description on the dataset page: https://huggingface.co/datasets/mattwesney/CoT_Temporal_Reasoning_Dataset.Quant-CoT-Factor-Reasoning-PreviewQuantitative Factor Generation: Chain-of-Thought (CoT) Trajectories
Dataset Description
This is a 100-episode preview of a proprietary Reinforcement Learning from Environment Feedback (RLEF) dataset. It is designed to fine-tune Large Language Models (LLMs) for institutional quantitative finance, specifically systematic factor discovery and vectorized Python execution.
The Architecture
The data captures multi-turn agentic loops where the LLM:
Formulates a cross-sectional equity factor… See the full description on the dataset page: https://huggingface.co/datasets/1Happy-neuron/Quant-CoT-Factor-Reasoning-Preview.ambari-reasoning-cot
