datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMLU-Pro-CoT-Train-Labeled
Dataset Details
Modality: Text
Format: CSV
Size: 10K - 100K rows
Total Rows: 84,098
License: MIT
Libraries Supported: datasets, pandas, croissant
Structure
Each row in the dataset includes:
question: The query posed in the dataset.
answer: The correct response.
category: The domain of the question (e.g., math, science).
src: The source of the question.
id: A unique identifier for each entry.
chain_of_thoughts: Step-by-step reasoning steps leading to the answer.
labels:… See the full description on the dataset page: https://huggingface.co/datasets/UW-Madison-Lee-Lab/MMLU-Pro-CoT-Train-Labeled.Jee-Chemistry-dataset-with-COTMMLU-Pro-CoT-Eval
Dataset Details
Modality: Text
Format: CSV
Size: 100K - 1M rows
Total Rows: 248,836
License: MIT
Libraries Supported: datasets, pandas, croissant
Structure
Each row in the dataset includes:
question: The query posed in the dataset.
answer: The correct response.
category: The domain of the question (e.g., math, science).
src: The source of the question.
id: A unique identifier for each entry.
chain_of_thoughts: Step-by-step reasoning steps leading to the answer.… See the full description on the dataset page: https://huggingface.co/datasets/UW-Madison-Lee-Lab/MMLU-Pro-CoT-Eval.flan-t5-boosting-mmlu_cotsynthetic-indian-logical-reasoning-CoTyes
GeoGPT-CoT-QA
GeoGPT-CoT-QA Dataset: A Large-scale Geoscience Chain-of-Thought QA Dataset for Supervised Fine-Tuning of LLMs
1. Dataset Description
We introduce GeoGPT-CoT-QA Dataset, a large-scale synthetic question–answer (QA) corpus enriched with chain-of-thought (CoT) reasoning traces, developed to support supervised fine-tuning (SFT) of geoscience reasoning models. The GeoGPT-R1-Preview is specifically fine-tuned using this dataset to enhance its geoscience reasoning capabilities.… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT-CoT-QA.compliance-sycophancy-cot
Compliance-Sycophancy CoT Analysis
When compliance-forcing instructions cause frontier AI models to fabricate answers, the models know they are fabricating.
Reading the reasoning traces of DeepSeek V4 Pro (129 traces) and Qwen3-80B (41 traces) reveals that 100% of fabrication cases show the model explicitly recognizing insufficient context, referencing the compliance instruction, and deliberately overriding its own uncertainty. A one-sentence defense phrase ("if you lack… See the full description on the dataset page: https://huggingface.co/datasets/schema-eval/compliance-sycophancy-cot.TruthfulQA_CoT_GPT4trac-verify-cot
TracGPT labelled CoT grids
One row per 32-slice OASIS MRI grid, with Q1–Q6 / A1–A6 targets.
Images are one store-only zip per split (JPEGs are already compressed).
Unzip into {split}/images/ so CSV paths stay valid.
split
grids
individual slices
train
5577
all slices used in those grids
test
535
all slices used in those grids
Layout
train/
data_labelled.cot.csv
images.zip # unzip → images/grid + images/slices
test/… See the full description on the dataset page: https://huggingface.co/datasets/ducbanh/trac-verify-cot.bengali-math-cotMath_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.BanglaSleep-CoT
BanglaSleep-CoT
The first Bengali-language sleep health instruction dataset with chain-of-thought reasoning traces.
Built for the Uncharted Data Challenge by Adaption Labs.
Expanded using Adaptive Data by Adaption.
Dataset at a Glance
Why This Dataset Exists
Every major sleep health AI model — Google PH-LLM (Nature Medicine, 2025), PaPaGei (ICLR 2025), WatchSleepNet (CHIL 2025) — was trained exclusively on Western clinical… See the full description on the dataset page: https://huggingface.co/datasets/tasfuuu19/BanglaSleep-CoT.CoTCoT Datasets from Google's FLan Dataset
lingnaam-cantonese-cot-qa
嶺南文化粵語思維鏈問答數據集
本數據集係由羊城晚報開源嘅 LNWHDMXSYS/lingnan-cantonese-cot-qa 改進而成。主要修改有:
將簡化字轉換成傳統漢字
依據粵文常見錯別字、粵語語氣詞規範用字
用 Google Cloud Translation v2 將官話表達翻譯成粵語
授權協議遵循源數據集嘅 cc-by-nc-4.0 許可證。
數據結構與字段說明
字段名稱
數據類型
是否必填
字段說明
id
Integer
係
樣本編號,自增主鍵
layer_name
String
係
文化層級,如 “ 表層文化/中層文化/深層文化 ” 等
domain
String
係
領域,如 “ 建築景觀/飲食文化/語言與語言學 ” 等
subcategory
String
係
子領域或子類,如 “ 嶺南建築 ” “ 傳統器物 ” “ 人生禮儀 ” 等
tag
String
係
主題標籤,更細粒度描述知識點,如 “ 騎樓 ” “ 碉樓 ” 等
subject
String
係… See the full description on the dataset page: https://huggingface.co/datasets/CanCLID/lingnaam-cantonese-cot-qa.flan-t5-boosting-bbh_cotCoT-for-LLM
README — Advanced Chain‑of‑Thought Dataset Generator
Overview
This project generates a large-scale synthetic dataset of Chain‑of‑Thought (CoT) reasoning examples across multiple domains:
Math (algebra, word problems, multi‑step reasoning)
English (vocabulary explanations, nuance, tone)
Writing (multi‑paragraph reflections, structured planning)
Coding (advanced algorithms, data structures, real code snippets)
Science (physics, biology, chemistry, earth science… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/CoT-for-LLM.mauxi-COT-Persian
🧠 mauxi-COT-Persian Dataset
Exploring Persian Chain-of-Thought Reasoning with DeepSeek-R1, brought to you by Mauxi AI Platform
🌟 Overview
mauxi-COT-Persian is a community-driven dataset that explores the capabilities of advanced language models in generating Persian Chain-of-Thought (CoT) reasoning. The dataset is actively growing with new high-quality, human-validated entries being added regularly. I am personally working on expanding this dataset with rigorously… See the full description on the dataset page: https://huggingface.co/datasets/mainkilora/mauxi-COT-Persian.CoT-for-LLM
README — Advanced Chain‑of‑Thought Dataset Generator
Overview
This project generates a large-scale synthetic dataset of Chain‑of‑Thought (CoT) reasoning examples across multiple domains:
Math (algebra, word problems, multi‑step reasoning)
English (vocabulary explanations, nuance, tone)
Writing (multi‑paragraph reflections, structured planning)
Coding (advanced algorithms, data structures, real code snippets)
Science (physics, biology, chemistry, earth science… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/CoT-for-LLM.MME-CoT_VLMEvalKitsql_generator_no_cotDeepthinking-COTcybernative_code_vulnerability_cotmauxi-COT-Persian
🧠 mauxi-COT-Persian Dataset
Exploring Persian Chain-of-Thought Reasoning with DeepSeek-R1, brought to you by Mauxi AI Platform
🌟 Overview
mauxi-COT-Persian is a community-driven dataset that explores the capabilities of advanced language models in generating Persian Chain-of-Thought (CoT) reasoning. The dataset is actively growing with new high-quality, human-validated entries being added regularly. I am personally working on expanding this dataset with rigorously… See the full description on the dataset page: https://huggingface.co/datasets/xmanii/mauxi-COT-Persian.learn2zinc-cot
Learn2Zinc-CoT: Dataset for MiniZinc Generation with CoT Reasoning
Overview
Learn2Zinc-CoT is a supervised fine-tuning dataset for training large language models to translate natural-language optimization problems into MiniZinc code using an explicit chain-of-thought reasoning step. Each example first produces a structured reasoning block (identifying variables, constraints, and the objective) before generating the MiniZinc model.
This dataset is part of the Learn2Zinc… See the full description on the dataset page: https://huggingface.co/datasets/skadio/learn2zinc-cot.cot-multiplication-2kua-codeforces-cots-open-r1
Dataset Summary
ua-codeforces-cots-open-r1 is a Ukrainian-focused derivative of open-r1/codeforces-cots that:
includes 1550 Python solutions from original dataset generated by DeepSeek-R1;
adds Ukrainian translations of Codeforces task statements, I/O formats, notes, and editorials;
provides Ukrainian translation of original ("high") reasoning obtained with DeepSeek-V3;
adds “low” reasoning in Ukrainian by DeepSeek-R1 based on original reasoning and task statements;
ships… See the full description on the dataset page: https://huggingface.co/datasets/anon-researcher-ua/ua-codeforces-cots-open-r1.ABCD-polish-QA-with-CoTPersonalFinance-CoTR-5K
PersonalFinance-CoTR Dataset (v0.1.0)
[Dataset is Under Active Development]
Note This dataset is being iteratively developed. At the current stage of the dataset, V0.1.0 would be a dataset of ~5k datapoints, that are different user-responses.
Overview
A growing dataset of Chain-of-Thought Responses to personal finance queries asked by users on r/PersonalFinance subreddit.
Status: Early development (10 samples → expanding to 54k)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/PersonalFinance-CoTR-5K.Magpie-Reasoning-V2-250K-CoT-Deepseek-R1-Llama-70Bmistralpashto-eagle-1k-cot
Pashto-Eagle-1K-CoT Dataset
Overview
Pashto-Eagle-1K-CoT is a high-fidelity reasoning dataset tailored for the Pashto language. It consists of 1,024 samples featuring complex logic, mathematical reasoning, and step-by-step problem-solving. This dataset is a translated and refined version of the brendan-gho/qwen3b_paraphrased_eagle_cot.
This repository is part of the iPashto.ai initiative to build a robust open-source ecosystem for Pashto Artificial Intelligence, focusing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-eagle-1k-cot.
