datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Fable-5-CoT-TracesPersonal collection of Fable 5 reasoning traces.
Filter out the decoy ones and you're good.
Have fun! (Also, star my repo https://github.com/FusionCube18712/claude-codex-auto-resume if you can)
Happy distilling.
DAG-MATH-Formatted-CoT
Benchmark Overview
This dataset card contains 2,894 gold-standard DAG-MATH formatted CoT from problems from Omni-MATH.
Top‑Level Schema
Each JSON file is a list with a single object describing the problem:
problem_id: integer identifier of the problem.
domain: list of strings describing the topic taxonomy.
difficulty: numeric difficulty indicator from 1 (easiest) to 6 (hardest).
problem_text: problem statement.
sample_id: sample identifier for the solution trace.… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/DAG-MATH-Formatted-CoT.X-CoT
X-CoT: Explainable Text-to-Video Retrieval Dataset
This repository contains the dataset for X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning.
This dataset expands existing text-to-video retrieval benchmarks with additional video annotations to support semantic understanding and reduce data bias. It is designed to facilitate explainable retrieval frameworks based on LLM Chain-of-Thought reasoning, aiming to improve retrieval performance and… See the full description on the dataset page: https://huggingface.co/datasets/prasannareddyp/X-CoT.pacman_hard_cot_chunk_k10_train
pacman_hard_cot_chunk_k10_train
BAGEL VLM-Gym world-model dataset (pacman / cot).
CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=10 steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching pacman checkpoint(s) under the companion model org; CoT and non-CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pacman_hard_cot_chunk_k10_train.model_20_tokens_10_specsdeepseek-v4-pro-math-cot-1k
DeepSeek V4 Pro Math CoT 1K
A small, high-signal supervised-fine-tuning (SFT) dataset of math reasoning traces. Problems were sampled from a Nemotron math problem set (originally sourced from StackExchange-Math and AoPS), answered by DeepSeek V4 Pro with thinking enabled at high reasoning effort, then independently reviewed by DeepSeek V4 Flash for correctness against the expected answer. Pathological reasoning traces (looping, run-away length, excessive Wait-style backtracking)… See the full description on the dataset page: https://huggingface.co/datasets/blythet/deepseek-v4-pro-math-cot-1k.pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot
PushT norm4 Visual Nomarker All-Step Thinking Trickiness COT
This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_norm4_visual_nomarker.
Each row contains one full successful trajectory from the first move through the final stop action.
Main files:
training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows.
testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows.
metadata/final_scan_validation.json: full local scan after repair… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot.deepseek-v4-flash-swe-cot
DeepSeek-V4-Flash SWE Agent Trajectories (with raw chain-of-thought)
795 multi-turn software-engineering agent trajectories generated by
DeepSeek-V4-Flash-0731 at reasoning_effort=max, each one executed in a real
repository inside an isolated container and verified by running the repository's own
tests. 469 are verified-correct.
Every assistant turn preserves reasoning_content — the model's raw chain-of-thought,
not a summary. That is the point of this dataset: the DeepSeek API… See the full description on the dataset page: https://huggingface.co/datasets/blythet/deepseek-v4-flash-swe-cot.novelist-cot-writing-raw-v1
Novelist: Human-Like Creative Writing Dataset (RAW)
This dataset is designed to train LLMs in high-quality creative writing. It focuses on narrative depth, coherent world-building, and logical character psychology.
The data was generated using DeepSeek-R1.
Dataset Overview
We focused on Quality over Quantity. The goal was to move away from generic "AI slop" and create text that feels grounded and intentional.
Total Tokens: ~29.4 Million
Total Examples: 3,369
Format:… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/novelist-cot-writing-raw-v1.Kukedlc__Qwen-2.5-7b-Spanish-o1-CoT-details
Dataset Card for Evaluation run of Kukedlc/Qwen-2.5-7b-Spanish-o1-CoT
Dataset automatically created during the evaluation run of model Kukedlc/Qwen-2.5-7b-Spanish-o1-CoT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Kukedlc__Qwen-2.5-7b-Spanish-o1-CoT-details.EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-details.Quazim0t0__CoT_Phi-details
Dataset Card for Evaluation run of Quazim0t0/CoT_Phi
Dataset automatically created during the evaluation run of model Quazim0t0/CoT_Phi
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Quazim0t0__CoT_Phi-details.sokoban_easy_cot_chunk_k1_train
sokoban_easy_cot_chunk_k1_train
BAGEL VLM-Gym world-model dataset (sokoban / cot).
CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=1 steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching sokoban checkpoint(s) under the companion model org; CoT and non-CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/sokoban_easy_cot_chunk_k1_train.pusht_cot_chunk_k3_train
pusht_cot_chunk_k3_train
BAGEL VLM-Gym world-model dataset (pusht / cot).
CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=3 steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching pusht checkpoint(s) under the companion model org; CoT and non-CoT
variants… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pusht_cot_chunk_k3_train.EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT-details.clean_cot_verification_340k元データ: https://huggingface.co/datasets/Zigeng/CoT-Verification-340k
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT-Verification-340k
データ件数: 140,980
平均トークン数: 602
最大トークン数: 2,040
合計トークン数: 84,894,510
ファイル形式: JSONL
ファイル分割数: 2
合計ファイルサイズ: 256.3 MB
加工内容:
データセットIDの付与: データフレームのインデックスに1を加算して、base_datasets_idとして新しいID列を付与しました。
response列のフィルタリング: response列が「Yes,」で始まる行のみを保持し、それ以外の行を除外しました。
prompt列の文字長によるフィルタリング: prompt列の文字列の長さが80… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/clean_cot_verification_340k.MATH_OOD_Test_D1_Base_Model_Eval_COTLLaVA-CoT-30k-jsonl-trainkitsphinxnautics-codeforces-cot-v3cot-oracle-eval-step-importance-thought-anchors
CoT Oracle Eval: step_importance_thought_anchors
Causal step importance identification from off-policy deepseek MATH rollouts. Source: uzaymacar/math-rollouts.
Part of the CoT Oracle Evals collection.
Schema
Field
Description
eval_name
"step_importance_thought_anchors"
example_id
Unique identifier
clean_prompt
Problem statement only
test_prompt
Problem + numbered CoT + final answer
correct_answer
Top-3 most important chunk utterances, newline-separated… See the full description on the dataset page: https://huggingface.co/datasets/japhba/cot-oracle-eval-step-importance-thought-anchors.cot-oracle-eval-thought-anchors
CoT Oracle Eval: step_importance_thought_anchors
Causal step importance identification from off-policy DeepSeek MATH rollouts.
Importance metric: resampling KL divergence (resampling_importance_kl) — measures the KL divergence of the answer distribution when a chunk is removed and the continuation is resampled (~100 rollouts per chunk). This is the standard counterfactual importance metric from the Thought Anchors paper, NOT importance++.
Source: uzaymacar/math-rollouts (Thought… See the full description on the dataset page: https://huggingface.co/datasets/japhba/cot-oracle-eval-thought-anchors.pusht_96_norm4_cot_chunk_k3_20260622_perseg
pusht_96_norm4_cot_chunk_k3_20260622_perseg
PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame re-grounding… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_cot_chunk_k3_20260622_perseg.maze2d_easy_native256_cot_chunk_kinf_20260707_perseg
maze2d_easy_native256_cot_chunk_kinf_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_kinf_20260707_perseg.cot-oracle-qwen3-8b-onpolicy-recipe
CoT Activation Oracle — On-Policy Qwen3-8B Training Recipe
A reproduction of the on-policy Qwen3-8B training mixture from
Building Better Activation Oracles
(Bauer, De Schamphelaere, Karvonen, Luick, Nanda).
This repository is a recipe card only — it documents the exact dataset
mixture, points at every source on the Hub, and gives regeneration instructions
for the pieces that are no longer available upstream. No third-party data is
re-hosted here; original datasets are linked… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/cot-oracle-qwen3-8b-onpolicy-recipe.clean_pubmedqa_mixtral_cot元データ: https://huggingface.co/datasets/HPAI-BSC/PubmedQA-Mixtral-CoT
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/PubmedQA-Mixtral-CoT
データ件数: 206,962
平均トークン数: 586
最大トークン数: 1,922
合計トークン数: 121,366,170
ファイル形式: JSONL
ファイル分割数: 3
合計ファイルサイズ: 532.2 MB
加工内容:
文字数によるフィルタリング:
question (質問) 列の文字数が 6,000文字を超える データを削除します。
response (応答) 列の文字数が 80,000文字を超える データを削除します。
応答 (response) の分割:
response 列を、思考プロセスを記述した「thought」部分と、最終的な結論である「answer」部分に分割します。
分割には Answer: や The answer… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/clean_pubmedqa_mixtral_cot.maze2d_easy_native256_cot_chunk_k5_20260707_perseg
maze2d_easy_native256_cot_chunk_k5_20260707_perseg
Maze2d (native 256px, JPEG q95; navigation with stop-required success, easy→hard split) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/maze2d_easy_native256_cot_chunk_k5_20260707_perseg.pusht_cot_chunk_kinf_train
pusht_cot_chunk_kinf_train
BAGEL VLM-Gym world-model dataset (pusht / cot).
CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=inf (open-loop; imagine the whole episode) steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching pusht checkpoint(s) under the… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pusht_cot_chunk_kinf_train.model_20_tokens_3_specspusht_96_int1_visual_nomarker_allstep_thinking_trickiness_cot
PushT int1 Visual Nomarker All-Step Thinking Trickiness COT
This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_int1_visual_nomarker.
Each row contains one full successful trajectory from the first move through the final stop action.
Main files:
training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows.
testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows.
Message format:
Each user turn is the PushT prompt text plus one… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_int1_visual_nomarker_allstep_thinking_trickiness_cot.maze2d_easy_cot_chunk_k3_train
maze2d_easy_cot_chunk_k3_train
BAGEL VLM-Gym world-model dataset (maze2d / cot).
CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=3 steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching maze2d checkpoint(s) under the companion model org; CoT and non-CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/maze2d_easy_cot_chunk_k3_train.
