datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.Domofon-Cot-Conversations-700k
Domofon-Cot-Conversations-700k
Synthetic XML conversation data for training small language models on reasoning,
instruction following, XML formatting, and tool-use traces.
Repository: domofon/Domofon-Cot-Conversations-700k
What is inside
The dataset contains cleaned generated XML conversations from six families:
conv: multi-turn factual conversations with tool-use traces.
instruct: text-processing instructions, including deterministic count tool calls.
ds:… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Domofon-Cot-Conversations-700k.cot-gemma4-26b-a4b
Gemma-4-26B-A4B-it Chain-of-Thought Oracle Corpus
Chain-of-thought rollouts generated with google/gemma-4-26B-A4B-it (MoE,
25.2B total / 3.8B active), in its native thinking mode, across a diverse suite
of reasoning tasks. Structure follows
ceselder/cot-oracle-corpus-v5
(CoT-only subset of the columns), built for chain-of-thought monitoring /
activation-oracle research.
2,121,354 rollouts over 212,161 unique problems (10 sampled
thinking rollouts per problem, temperature 0.8).… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-gemma4-26b-a4b.measuring_cot_monitorability_transcripts
Measuring Chain-of-Thought Monitorability Transcripts
This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness.
We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.compliance-sycophancy-cot
Compliance-Sycophancy CoT Analysis
When compliance-forcing instructions cause frontier AI models to fabricate answers, the models know they are fabricating.
Reading the reasoning traces of DeepSeek V4 Pro (129 traces) and Qwen3-80B (41 traces) reveals that 100% of fabrication cases show the model explicitly recognizing insufficient context, referencing the compliance instruction, and deliberately overriding its own uncertainty. A one-sentence defense phrase ("if you lack… See the full description on the dataset page: https://huggingface.co/datasets/schema-eval/compliance-sycophancy-cot.deepseek-v4-pro-math-cot-1k
DeepSeek V4 Pro Math CoT 1K
A small, high-signal supervised-fine-tuning (SFT) dataset of math reasoning traces. Problems were sampled from a Nemotron math problem set (originally sourced from StackExchange-Math and AoPS), answered by DeepSeek V4 Pro with thinking enabled at high reasoning effort, then independently reviewed by DeepSeek V4 Flash for correctness against the expected answer. Pathological reasoning traces (looping, run-away length, excessive Wait-style backtracking)… See the full description on the dataset page: https://huggingface.co/datasets/blythet/deepseek-v4-pro-math-cot-1k.Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.BIRD-Verified-CoT-2462-GPT5.4
BIRD-Verified-CoT-2462 (GPT-5.4 distilled)
Likely the first publicly available CoT-augmented Text-to-SQL dataset built on top of expert-verified BIRD data.
This dataset combines two state-of-the-art ingredients:
ReViSQL's BIRD-Verified subset — 2,462 SQL-expert verified examples (multi-round review by UIUC team), eliminating the ~50% annotation noise of the original BIRD train set.
GPT-5.4 (via Codex CLI) — distilled into structured 6-section Chain-of-Thought traces using… See the full description on the dataset page: https://huggingface.co/datasets/wenyupapa/BIRD-Verified-CoT-2462-GPT5.4.BanglaSleep-CoT
BanglaSleep-CoT
The first Bengali-language sleep health instruction dataset with chain-of-thought reasoning traces.
Built for the Uncharted Data Challenge by Adaption Labs.
Expanded using Adaptive Data by Adaption.
Dataset at a Glance
Why This Dataset Exists
Every major sleep health AI model — Google PH-LLM (Nature Medicine, 2025), PaPaGei (ICLR 2025), WatchSleepNet (CHIL 2025) — was trained exclusively on Western clinical… See the full description on the dataset page: https://huggingface.co/datasets/tasfuuu19/BanglaSleep-CoT.China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740… See the full description on the dataset page: https://huggingface.co/datasets/a13905873166/China-K12-STEM-10K-CoT-Reasoning.clean_cot_verification_340k元データ: https://huggingface.co/datasets/Zigeng/CoT-Verification-340k
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT-Verification-340k
データ件数: 140,980
平均トークン数: 602
最大トークン数: 2,040
合計トークン数: 84,894,510
ファイル形式: JSONL
ファイル分割数: 2
合計ファイルサイズ: 256.3 MB
加工内容:
データセットIDの付与: データフレームのインデックスに1を加算して、base_datasets_idとして新しいID列を付与しました。
response列のフィルタリング: response列が「Yes,」で始まる行のみを保持し、それ以外の行を除外しました。
prompt列の文字長によるフィルタリング: prompt列の文字列の長さが80… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/clean_cot_verification_340k.cot-qa-gemma4-26b-a4b
cot-qa-gemma4-26b-a4b — Activation-Oracle Probes
Probing questions over cds-jb/gemma4-26b-a4b-cot-oracle-corpus
(chain-of-thought rollouts from google/gemma-4-26B-A4B-it). Each row is ONE
probe: a question about a gemma-4 CoT that is hard-from-text but
easy-from-the-latent-activation, for evaluating an activation-oracle M.
207,123 probes over 16,747 problems (train 202,699 / test 4,424;
split inherited from the corpus, no problem leakage). Generated by
claude-sonnet-4-6 via the… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-qa-gemma4-26b-a4b.clean_pubmedqa_mixtral_cot元データ: https://huggingface.co/datasets/HPAI-BSC/PubmedQA-Mixtral-CoT
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/PubmedQA-Mixtral-CoT
データ件数: 206,962
平均トークン数: 586
最大トークン数: 1,922
合計トークン数: 121,366,170
ファイル形式: JSONL
ファイル分割数: 3
合計ファイルサイズ: 532.2 MB
加工内容:
文字数によるフィルタリング:
question (質問) 列の文字数が 6,000文字を超える データを削除します。
response (応答) 列の文字数が 80,000文字を超える データを削除します。
応答 (response) の分割:
response 列を、思考プロセスを記述した「thought」部分と、最終的な結論である「answer」部分に分割します。
分割には Answer: や The answer… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/clean_pubmedqa_mixtral_cot.cleand_HangHor_FinQA_CoT_Small元データ: https://huggingface.co/datasets/HangHor/FinQA_CoT_Small
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/FinQA_CoT_Small
データ件数: 309
平均トークン数: 1,036
最大トークン数: 2,787
合計トークン数: 320,198
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 1.3 MB
加工内容:
1. データセットの準備と初期設定
データソース: Hugging Face Hub上の HangHor/FinQA_CoT_Small データセットを読み込んでいます。
トークナイザー: トークン数の計算には deepseek-ai/DeepSeek-R1-Distill-Qwen-32B モデルのトークナイザーが使用されています。
列の役割設定: 元のデータセットの Question、Context、Answer… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_HangHor_FinQA_CoT_Small.when-the-cot-knows-better
When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models
Accepted at the ICML 2026 Workshop on Failure Modes in Agentic AI (FAGEN)
This repository contains the dataset for the paper When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models.
Dataset Summary
This dataset contains the evaluation artifacts for the paper "When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models". It… See the full description on the dataset page: https://huggingface.co/datasets/UVSKKR/when-the-cot-knows-better.cleand_moremilk_CoT_Reasoning_Quantom_Physics_And_Computing元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Quantom_Physics_And_Computing
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Quantom_Physics_And_Computing
データ件数: 2,862
平均トークン数: 1,110
最大トークン数: 2,334
合計トークン数: 3,175,666
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 15.5 MB
加工内容:
メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。
難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Quantom_Physics_And_Computing.cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Scientific_Discovery_and_Research
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Scientific_Discovery_and_Research
データ件数: 3,733
平均トークン数: 1,193
最大トークン数: 2,489
合計トークン数: 4,453,517
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 23.2 MB
加工内容:
メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。
難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research.cleand_cw18_lean-six-sigma-cot-500元データ: https://huggingface.co/datasets/cw18/lean-six-sigma-cot-500
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/lean-six-sigma-cot-500
データ件数: 215
平均トークン数: 514
最大トークン数: 591
合計トークン数: 110,520
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 602.9 KB
加工内容:
文字列長によるフィルタリング:
instruction列(質問)の文字数が6000文字を超える行を除外しました。
output列(思考)の文字数が80000文字を超える行を除外しました。
思考タグの除去と分割:
IS_THINKTAGがFalseに設定されているため、output列をSPLIT_KEYWORD (**Final Toolset… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_cw18_lean-six-sigma-cot-500.cleand_open-r1_codeforces-cots元データ: https://huggingface.co/datasets/open-r1/codeforces-cots
データ件数: 5,334
平均トークン数: 11512
最大トークン数: 30,725
合計トークン数: 61,406,976
ファイル形式: JSONL
ファイルサイズ: 213.1 MB
加工内容
solutions_w_editorials_decontaminatedを使用
停止理由をstopに限定
トークン処理が重たいので、文字数でフィルター
prompt < 6000
generation < 80000
accepted_solutionsがあるもの
thinkタグ除去
繰り返し除去
