datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NuminaMath-CoT
Dataset Card for NuminaMath CoT
Dataset Summary
Approximately 860k math problems, where each solution is formatted in a Chain of Thought (CoT) manner. The sources of the dataset range from Chinese high school math exercises to US and international mathematics olympiad competition problems. The data were primarily collected from online exam paper PDFs and mathematics discussion forums. The processing steps include (a) OCR from the original PDFs, (b) segmentation… See the full description on the dataset page: https://huggingface.co/datasets/AI-MO/NuminaMath-CoT.Alpaca-CoT
Instruction-Finetuning Dataset Collection (Alpaca-CoT)
This repository will continuously collect various instruction tuning datasets. And we standardize different datasets into the same format, which can be directly loaded by the code of Alpaca model.
We also have conducted empirical study on various instruction-tuning datasets based on the Alpaca model, as shown in https://github.com/PhoebusSi/alpaca-CoT.
If you think this dataset collection is helpful to you, please like… See the full description on the dataset page: https://huggingface.co/datasets/QingyiSi/Alpaca-CoT.cot-eval-traces-2.0CoTracker3_Kubric
Kubric Dataset for CoTracker 3
Overview
This dataset was specifically created for training CoTracker 3, a state-of-the-art point tracking model. The dataset was generated using the Kubric engine.
Dataset Specifications
Size: ~6,000 sequences
Resolution: 512×512 pixels
Sequence Length: 120 frames per sequence
Camera Movement: Carefully rendered with subtle camera motion to simulate realistic scenarios
Format: Generated using Kubric engine
Usage
The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/CoTracker3_Kubric.Zebra-CoT
Zebra‑CoT
A diverse large-scale dataset for interleaved vision‑language reasoning traces.
Dataset Description
Zebra‑CoT is a diverse large‑scale dataset with 182,384 samples containing logically coherent interleaved text‑image reasoning traces across four major categories: scientific reasoning, 2D visual reasoning, 3D visual reasoning, and visual logic & strategic games.
Dataset Structure
Each example in Zebra‑CoT consists of:
Problem statement:… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT.Visual-CoT
VisCoT Dataset Card
There is a shortage of multimodal datasets for training multi-modal large language models (MLLMs) that require to identify specific regions in an image for additional attention to improve response performance. This type of dataset with grounding bbox annotations could possibly help the MLLM output intermediate interpretable attention area and enhance performance.
To fill the gap, we curate a visual CoT dataset. This dataset specifically focuses on identifying… See the full description on the dataset page: https://huggingface.co/datasets/deepcs233/Visual-CoT.codeforces-cots
Dataset Card for CodeForces-CoTs
Dataset description
CodeForces-CoTs is a large-scale dataset for training reasoning models on competitive programming tasks. It consists of 10k CodeForces problems with up to five reasoning traces generated by DeepSeek R1. We did not filter the traces for correctness, but found that around 84% of the Python ones pass the public tests.
The dataset consists of several subsets:
solutions: we prompt R1 to solve the problem and produce code.… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/codeforces-cots.efficient-cot
To be written.
OpenMath-Vision-CoT-10kScaffold-CoT
Scaffold-CoT
A ~3.8M example, ~3B token CoT dataset designed around helping small models think more
concisely, accurately and reliably. Every example is labelled with a domain and a subdomain, so
you can train on exactly the slice you want.
Why this exists
When using small models (Under 5B parameters), I noticed freeform CoT does not really add much
in terms of capability, and usually results in more confusing, poorly structured and inaccurate
responses.
This… See the full description on the dataset page: https://huggingface.co/datasets/Specific-Labs/Scaffold-CoT.Audio-Reasoner-CoTAMMLU-Pro-CoT-Train-43KStep-CoT
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
Dataset Summary
Step-CoT is a large-scale medical reasoning dataset designed to improve interpretability in Medical Visual Question Answering (Med-VQA). It contains over 10,000 real clinical chest X-ray cases and 70,000 VQA pairs, each structured into a seven-step diagnostic workflow that mirrors clinical reasoning:
Abnormal Radiodensity Detection
Lesion Distribution
Radiographic Pattern… See the full description on the dataset page: https://huggingface.co/datasets/fl-15o/Step-CoT.ISCSLP2026-CoT-TTS
ISCSLP 2026 CoT-TTS Dataset
Dataset Overview
This dataset is prepared for the ISCSLP 2026 CoT-TTS Challenge and is designed to support research on context-aware, expressive, and CoT-guided speech generation. It is constructed from speech-rich media sources, including films, TV dramas, radio dramas, and short dramas, where dialogue often contains rich conversational context, speaker interactions, scene changes, and emotional variation. Each sample is organized… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/ISCSLP2026-CoT-TTS.libero_cotThis dataset was created using LeRobot.
It contains embodied Chain-of-Thought (CoT) demonstrations for the LIBERO benchmark, featuring paired reasoning and action traces. It was curated as part of the DeepThinkVLA project using a two-stage data engine that distills key frames with a cloud LVLM and scales to full trajectories via a fine-tuned local VLM.
Paper: DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
Repository: https://github.com/OpenBMB/DeepThinkVLA… See the full description on the dataset page: https://huggingface.co/datasets/yinchenghust/libero_cot.Taur_CoT_Analysis_Project___gpt-4o-2024-08-06mmlu_cotCottonWeedDet12
Dataset Card for CottonWeedDet12
CottonWeedDet12 is a 12-class weed object-detection dataset for cotton production systems in the southern U.S., consisting of 5,648 RGB field images with 9,370 bounding box annotations collected in Michigan State University MEFAS field trials during 2021-2022. It is the companion dataset for the YOLOWeeds benchmark of YOLO object detectors.
This is a FiftyOne dataset with 5648 samples.
Installation
If you haven't already, install… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/CottonWeedDet12.bbh-cotReason-RFT-CoT-Dataset
🤗 Reason-RFT CoT Dateset
The full dataset used in our project "Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning".
⭐️ Project │ 🌎 Github │ 🔥 Models │ 📑 ArXiv │ 💬 WeChat
🤖 RoboBrain: Aim to Explore ReasonRFT Paradigm to Enhance RoboBrain's Embodied Reasoning Capabilities.
♣️ Quick Start
Please refer to Dataset Preparation
🔥 Overview
Visual reasoning abilities play a crucial role in understanding complex multimodal… See the full description on the dataset page: https://huggingface.co/datasets/tanhuajie2001/Reason-RFT-CoT-Dataset.exp11-cot-leakage
Chain-of-Thought Leakage RL
TL;DR
What we're doing. Online GRPO on a 49B model (Tim Hua's "Wood organism" — a Nemotron Super 49B fine-tuned to behave differently when it thinks it's being evaluated by Wood Labs). The reward model is gpt-oss-120b reading Anthropic's Constitution as a system prompt, with the model's chain-of-thought visible. Each rollout gets two ratings — one with CoT shown to the judge, one with only the post-</think> response. The reward we train… See the full description on the dataset page: https://huggingface.co/datasets/rlundqvist/exp11-cot-leakage.APIGen-MT-5k-with-cot-v1-deepseek_deepseekCoT-Collection"""
_LICENSE = "CC BY 4.0"
_HOMEPAGE = "https://github.com/kaistAI/CoT-Collection"
_LANGUAGES = {
"en": "English",
}
# _ALL_LANGUAGES = "all_languages"
class CoTCollectionMultiConfig(datasets.BuilderConfig):embodied_features_and_demos_liberoDataset for Embodied Chain-of-Thought Reasoning for LIBERO-90, as used by ECoT-Lite.
TFDS Demonstration Data
The TFDS dataset contains successful demonstration trajectories for LIBERO-90 (50 trajectories for each of 90 tasks). It was created by rolling out the actions provided in the original LIBERO release and filtering out all unsuccessful ones, leaving 3917 successful demo trajectories. This is done via a modified version of a script from the MiniVLA codebase. In addition to… See the full description on the dataset page: https://huggingface.co/datasets/Embodied-CoT/embodied_features_and_demos_libero.mmlu-pro-self-cot-deepseek-r1Use deepseek-r1 to generate COT in few-shot examples.
bert_cot_em
Can you tell a model is about to misbehave by reading its reasoning?
Short answer: no — but you can change what it does by writing its reasoning for it.
This repo studies a large language model that has been deliberately made
misaligned, and asks whether its chain-of-thought (the "thinking out loud" it
does before answering) gives away that a harmful answer is coming.
The setup in plain terms
Researchers found that fine-tuning a model on bad medical advice makes… See the full description on the dataset page: https://huggingface.co/datasets/mild-rgb/bert_cot_em.LLaVA-CoT-100k
Dataset Card for LLaVA-CoT
The LLaVA-CoT-100k dataset is introduced in the paper LLaVA-CoT: Let Vision Language Models Reason Step-by-Step. This dataset is designed to enable Vision-Language Models (VLMs) to perform autonomous multistage reasoning, integrating samples from various visual question-answering sources with structured reasoning annotations. It aims to address the challenges VLMs face in systematic and structured reasoning for complex visual question-answering tasks.… See the full description on the dataset page: https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k.Taur_CoT_Analysis_Project___meta-llama__Meta-Llama-3.1-8B-InstructMath-CoT-44k-Qwen3-32b-n32-16384-with-logprob-and-entropy
Qwen3-32B Math n32 16384 (44k Queries)
This dataset contains multi-sampled rollout traces from Qwen3-32B on around 44k math queries.
For each query, the model is rolled out 32 times with a maximum generation length of 16384 tokens.
Each response is annotated with answer correctness (acc_reward), and includes token-level statistics (action_entropy, action_log_probs) for further analysis and research.
Resources
Paper: Rethinking Generalization in Reasoning SFT: A… See the full description on the dataset page: https://huggingface.co/datasets/jasonrqh/Math-CoT-44k-Qwen3-32b-n32-16384-with-logprob-and-entropy.NuminaMath-QwQ-CoT-5M
INTELLECT-MATH: Frontier Mathematical Reasoning through Better Initializations for Reinforcement Learning
INTELLECT-MATH is a 7B parameter model optimized for mathematical reasoning. It was trained in two stages, an SFT stage, in which the model was fine-tuned on verified QwQ outputs, and an RL stage, in which the model was trained using the PRIME-RL recipe.
We demonstrate that the quality of our SFT data can impact the performance and training speed of the RL stage: Due to its… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/NuminaMath-QwQ-CoT-5M.
