datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
compliance-sycophancy-cot
Compliance-Sycophancy CoT Analysis
When compliance-forcing instructions cause frontier AI models to fabricate answers, the models know they are fabricating.
Reading the reasoning traces of DeepSeek V4 Pro (129 traces) and Qwen3-80B (41 traces) reveals that 100% of fabrication cases show the model explicitly recognizing insufficient context, referencing the compliance instruction, and deliberately overriding its own uncertainty. A one-sentence defense phrase ("if you lack… See the full description on the dataset page: https://huggingface.co/datasets/schema-eval/compliance-sycophancy-cot.Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.BanglaSleep-CoT
BanglaSleep-CoT
The first Bengali-language sleep health instruction dataset with chain-of-thought reasoning traces.
Built for the Uncharted Data Challenge by Adaption Labs.
Expanded using Adaptive Data by Adaption.
Dataset at a Glance
Why This Dataset Exists
Every major sleep health AI model — Google PH-LLM (Nature Medicine, 2025), PaPaGei (ICLR 2025), WatchSleepNet (CHIL 2025) — was trained exclusively on Western clinical… See the full description on the dataset page: https://huggingface.co/datasets/tasfuuu19/BanglaSleep-CoT.lingnaam-cantonese-cot-qa
嶺南文化粵語思維鏈問答數據集
本數據集係由羊城晚報開源嘅 LNWHDMXSYS/lingnan-cantonese-cot-qa 改進而成。主要修改有:
將簡化字轉換成傳統漢字
依據粵文常見錯別字、粵語語氣詞規範用字
用 Google Cloud Translation v2 將官話表達翻譯成粵語
授權協議遵循源數據集嘅 cc-by-nc-4.0 許可證。
數據結構與字段說明
字段名稱
數據類型
是否必填
字段說明
id
Integer
係
樣本編號,自增主鍵
layer_name
String
係
文化層級,如 “ 表層文化/中層文化/深層文化 ” 等
domain
String
係
領域,如 “ 建築景觀/飲食文化/語言與語言學 ” 等
subcategory
String
係
子領域或子類,如 “ 嶺南建築 ” “ 傳統器物 ” “ 人生禮儀 ” 等
tag
String
係
主題標籤,更細粒度描述知識點,如 “ 騎樓 ” “ 碉樓 ” 等
subject
String
係… See the full description on the dataset page: https://huggingface.co/datasets/CanCLID/lingnaam-cantonese-cot-qa.mauxi-COT-Persian
🧠 mauxi-COT-Persian Dataset
Exploring Persian Chain-of-Thought Reasoning with DeepSeek-R1, brought to you by Mauxi AI Platform
🌟 Overview
mauxi-COT-Persian is a community-driven dataset that explores the capabilities of advanced language models in generating Persian Chain-of-Thought (CoT) reasoning. The dataset is actively growing with new high-quality, human-validated entries being added regularly. I am personally working on expanding this dataset with rigorously… See the full description on the dataset page: https://huggingface.co/datasets/mainkilora/mauxi-COT-Persian.mauxi-COT-Persian
🧠 mauxi-COT-Persian Dataset
Exploring Persian Chain-of-Thought Reasoning with DeepSeek-R1, brought to you by Mauxi AI Platform
🌟 Overview
mauxi-COT-Persian is a community-driven dataset that explores the capabilities of advanced language models in generating Persian Chain-of-Thought (CoT) reasoning. The dataset is actively growing with new high-quality, human-validated entries being added regularly. I am personally working on expanding this dataset with rigorously… See the full description on the dataset page: https://huggingface.co/datasets/xmanii/mauxi-COT-Persian.PersonalFinance-CoTR-5K
PersonalFinance-CoTR Dataset (v0.1.0)
[Dataset is Under Active Development]
Note This dataset is being iteratively developed. At the current stage of the dataset, V0.1.0 would be a dataset of ~5k datapoints, that are different user-responses.
Overview
A growing dataset of Chain-of-Thought Responses to personal finance queries asked by users on r/PersonalFinance subreddit.
Status: Early development (10 samples → expanding to 54k)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/PersonalFinance-CoTR-5K.pashto-eagle-1k-cot
Pashto-Eagle-1K-CoT Dataset
Overview
Pashto-Eagle-1K-CoT is a high-fidelity reasoning dataset tailored for the Pashto language. It consists of 1,024 samples featuring complex logic, mathematical reasoning, and step-by-step problem-solving. This dataset is a translated and refined version of the brendan-gho/qwen3b_paraphrased_eagle_cot.
This repository is part of the iPashto.ai initiative to build a robust open-source ecosystem for Pashto Artificial Intelligence, focusing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-eagle-1k-cot.flanv2_cot_dedepulicated
FLAN v2 Cot Deduplicated Dataset
Data Preprocessing
Remove instructions with less than 100 tokens in 'targets'.
Dedepulicate Dataset using cosine similarity with a threshold of 0.95.
Code
Github repo : https://github.com/AJlearner46/Deduplicate-flanv2-finetune-LLaMa3-
Acknowledgments
The original dataset is provided by SirNeural/flan_v2.
Tokenizer used: bert-base-uncased from Hugging Face.
pashto-otter-cot
Pashto-Otter-CoT Dataset
Overview
Pashto-Otter-CoT is a first-of-its-kind dataset specifically designed to bring Chain-of-Thought (CoT) Reasoning capabilities to Pashto language models. This dataset is a translated and curated version of a subset of the brendan-gho/gemma4b_paraphrased_otter_cot.
This project is part of the iPashto.ai initiative, led by Nassim الله (nassimjp), aimed at creating high-quality linguistic resources for the Pashto language.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-otter-cot.pashto-dragon-1k-cot
Pashto-Dragon-1K-CoT Dataset
Overview
Pashto-Dragon-1K-CoT is a specialized reasoning dataset containing 1,000+ samples, meticulously translated into Pashto to facilitate the development of advanced Chain-of-Thought (CoT) capabilities in Pashto LLMs. This dataset is a high-quality derivative of the brendan-gho/qwen3b_paraphrased_dragon_cot.
This repository is a core component of the iPashto.ai mission to move beyond simple web-scraping and focus on "Verified Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-dragon-1k-cot.pashto-qwen-1k-cot
Pashto-Qwen-1K-CoT Dataset
Overview
Pashto-Qwen-1K-CoT is a high-quality reasoning dataset consisting of 1,024 samples, specifically curated to enhance the Chain-of-Thought (CoT) capabilities of Pashto language models. This dataset is a translated version of a subset from brendan-gho/qwen3b_paraphrased_cat_cot.
By focusing on "Reasoning" rather than just "Information," this dataset helps models like Baran and Roshan develop logical thinking paths in the Pashto language.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-qwen-1k-cot.schema_cot_reasoning
⊙ Prompt Programs for Agentic Reasoning
Programmable task-dependent COTs for agentic reasoning.
A 100-row seed dataset for programmable cognition.
Each row defines:
Prompt template + input binding + explanation + task-dependent reasoning program
Pipeline:
intake → binding → procedure → output
Schema
Column
Meaning
ID
Stable row ID
Name
Task name
Prompt
Prompt template using {{VARIABLE}}
Expression
Input binding using $.path
Explanation
Binding… See the full description on the dataset page: https://huggingface.co/datasets/bitwikiorg/schema_cot_reasoning.
