datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nemotron-student-fail-v41-clean-thinking
Nemotron-fail / DeepSeek-V4.1 clean and action-only trajectories
DeepSeek-V4.1 reward-1 trajectories for tasks on which the Nemotron student
did not obtain reward 1. This release was rebuilt from the complete reward-1
audit under v54-high-precision-canonical-reconstruction-relations.
Training paths
Path
Rows
Unique tasks
Thinking
Use
data/strict/train.jsonl.gz
12
12
Preserved and clean
Raw-thinking SFT
data/hybrid/train.jsonl.gz
58
58
Only… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuanhucs/nemotron-student-fail-v41-clean-thinking.agent-think-tool_use
Agent Think Tool Use
Датасет многошаговых агентных сессий для дообучения моделей работе с кодом, инструментами и инженерными задачами. Записи содержат пользовательские требования, комментарии агента во время работы, decision summaries, вызовы инструментов, результаты запусков, обработку ошибок и финальную проверку.
Каждый shard представляет отдельную связанную сессию, а не отдельный вопрос и ответ. Данные охватывают исследование задачи, работу с документацией, проектирование… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/agent-think-tool_use.Dolci-Think-SFT-translated
Dolci-Think-SFT-translated
Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations.
Columns
Each row is a translated conversation plus the result of a post-translation quality filter:
id — source record id.
messages — the translated conversation (list of {content, role}).
filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.thinking-rollouts
thinking-rollouts
Unconstrained rollouts from thinking (chain-of-thought) models on DS-1000 and LiveCodeBench, CoT
saved verbatim alongside the final answer. Format per genlm/rollouts issue #5; schema is a superset
of temperature-sweep-data.
Hive-partitioned Parquet, thinking_mode folded into the model tag:
rollouts/domain=<dataset>/model=<tag>/temp=<temp>/data.parquet (tags like qwen3-8b-think,
qwen3-1.7b-nothink). 100 samples/instance.
Columns: model, thinking_mode, temp… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/thinking-rollouts.openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 5.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 1.3000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 4.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.0.5M-thinking
0.5M Thinking Dataset
This dataset contains responses generated by MiniMax-M2.1 for user questions from the a-m-team/AM-DeepSeek-R1-Distilled-1.4M dataset (am_0.5M subset).
Dataset Description
The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation.
Metric
Value
Examples
499,157
Total Tokens
3,732,749,397
Avg Tokens/Example
7,478
Source Dataset… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/0.5M-thinking.thinking-cap-tier-raw-traces
Thinking Cap Tier Raw Traces (TCS v4)
[!IMPORTANT]
Dataset Release v1.2 (Sept 2026) — Clean Delimiters & Zero-Padding Architecture:
All 38,158 candidate reasoning traces across all 4 tiers (candidates_low.jsonl, candidates_mid.jsonl, candidates_high.jsonl, candidates_xhigh.jsonl) are 100% sanitized:
Zero batch-padding residues (<|pad|>): Completely purged across all records.
Strict Delimiter Integrity: Generation blocks cleanly separate thought deliberation tags… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/thinking-cap-tier-raw-traces.scale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance3393-trajectories7928
Scale-SWE DeepSeek V4 Flash 0731 Think Rollouts
Successful AweAgent trajectories generated with deepseek-v4-flash-0731 in think mode.
Dataset summary
Source task instances: 3,393
Rollouts per source instance: 4
Total attempted rollouts: 13,572
Successful exported trajectories: 7,928
Unique instances represented by successful trajectories: 2,250
Scaffold: aweagent
Tool-call format: openai_function
The export retains assistant reasoning_content, function tool… See the full description on the dataset page: https://huggingface.co/datasets/wjn922-01/scale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance3393-trajectories7928.Soofi-Think-SFT-V2-firsthalf-DE
Soofi-Think-SFT-V2-firsthalf-DE
German-translated version of toroe/Soofi-Think-SFT-V2-firsthalf — a large-scale supervised fine-tuning dataset featuring chain-of-thought reasoning traces (<think>...</think>) across math, science, code, and general instruction-following tasks.
The translation was produced using Qwen3-32B via vLLM, applying professional-grade translation prompts with formal German register (Sie-form for professional/technical content).
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/toroe/Soofi-Think-SFT-V2-firsthalf-DE.Soofi-Think-SFT-V2-firsthalf-FR
Soofi-Think-SFT-V2-firsthalf-FR
French-translated version of toroe/Soofi-Think-SFT-V2-firsthalf — a large-scale supervised fine-tuning dataset featuring chain-of-thought reasoning traces (<think>...</think>) across math, science, code, tool-calling, and general instruction-following tasks.
The translation was produced using Qwen3-32B via vLLM, applying professional-grade translation prompts targeting standard French suitable for international francophone audiences.… See the full description on the dataset page: https://huggingface.co/datasets/toroe/Soofi-Think-SFT-V2-firsthalf-FR.arxiv-qa-thinking
ArXiv Q&A with Thinking Dataset
This dataset contains question-answer pairs generated by MiniMax-M2.1 based on academic articles from PursuitOfDataScience/arxiv-llama4-maverick-abstract.
Dataset Description
For each academic article, the model generates:
Thinking process: The model's reasoning wrapped in <think> tags
Question: An insightful question testing understanding of key concepts
Answer: A detailed answer based on the article content
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/arxiv-qa-thinking.Chinese-Qwen3-235B-Thinking-2507-Distill-100k
📌 Note: The English translation of this dataset card is provided below.
Chinese-Qwen3-235B-Thinking-2507-Distill-100k
Dataset Summary
Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。
该数据集覆盖了多个重要领域:
数学与工程任务(Mathematics, Applied Math, Advanced Math)
通用知识与写作(General Knowledge, Language & Writing)
技术与编程(Technology & Programming)
商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.0.9M-thinking
0.9M Thinking Dataset
This dataset contains responses generated by MiniMax-M2.1 for user questions from the a-m-team/AM-DeepSeek-R1-Distilled-1.4M dataset (am_0.9M subset).
Dataset Description
The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation.
Metric
Value
Examples
897,522
Total Tokens
5,954,272,687
Avg Tokens/Example
6,634
Source Dataset… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/0.9M-thinking.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 2.2000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think.gsm8k-thinking
GSM8K Thinking
This dataset contains responses generated by MiniMax-M2.1 for math word problems from the openai/gsm8k dataset.
Dataset Description
The dataset captures both the extended thinking process and final answers from MiniMax-M2.1, with reasoning wrapped in <think> tags for easy separation.
Metric
Value
Train Examples
7,473
Test Examples
1,319
Total Examples
8,792
Total Tokens
10,506,774
Avg Tokens/Example
1,195
Source Dataset… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/gsm8k-thinking.ru-reasoning_effort-sft_dpo_think_gpt
NotEvilAI/ru-reasoning_effort-sft_dpo_think_gpt
NotEvilAI/ru-reasoning_effort-sft_dpo_think_gpt -
синтетический датасет для поддержки генерации ризонинга на русском языке с вариативным объёмом thinking(reasoning_effort).
Reasoning_effort представлен в виде системного промта Reasoning: [effort], где effort - одно из следующих значений:
low, medium, high - стандартные значения минимального, среднего и большого ризонинга для gpt-oss-20b/gpt-oss-120b
none - отключить ризонинг, в… See the full description on the dataset page: https://huggingface.co/datasets/NotEvilAI/ru-reasoning_effort-sft_dpo_think_gpt.olmo-3-7b-think_ifeval
allenai/OLMo-3-7B-Think — ifeval
Model outputs from the micro-creativity inference suite.
Model: allenai/OLMo-3-7B-Think
Dataset: ifeval (541 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)
raw_output
Full… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/olmo-3-7b-think_ifeval.harvey-eval-gpt56sol-qwen35-9b-base-20t-think
harvey-eval-gpt56sol-qwen35-9b-base-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 3.9000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent runs. These are not new… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-base-20t-think.MBPP-Thinking-Gate-1k
MBPP Thinking-Gate SFT Dataset
This package contains two related assets:
Ready 1,000-row MBPP-style dataset (all.jsonl, train.jsonl, validation.jsonl).
It is synthetic and designed to test/train autonomous routing between <DIRECT> and <THINK>.
Official-MBPP builder (build_from_official_mbpp.py).
Run this to create the production dataset from the official Google Research MBPP source.
Why two response modes?
The training target starts with one of two routing… See the full description on the dataset page: https://huggingface.co/datasets/islam-kamel/MBPP-Thinking-Gate-1k.Qwen3.8-27B-thinking-completions
Qwen3.8-27B thinking-mode completions
17,022 prompts from public chat, math and code datasets, each answered once by Qwen3.8-27B (FP8 checkpoint) in thinking mode with its recommended sampling settings (54M completion tokens). Every sample has the reasoning trace and the final answer, as text and as the exact token ids.
The set was generated to train speculative-decoding drafters for this model, so it records the model's own sampled distribution rather than greedy output or… See the full description on the dataset page: https://huggingface.co/datasets/JonasLoos/Qwen3.8-27B-thinking-completions.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 3.1000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 3.4000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think.wmrl-v4-scale70-3m-agentic-eval-20t-think
scale70-3m Agentic evaluation, 20 turns, thinking enabled
1,000 saved evaluation attempts: 250 C&H firm-knowledge tasks × four samples.
The base model is Qwen/Qwen3.5-9B at revision c202236235762e1c871ad0ccb60c8ee5ba337b9a, trained for two
epochs on a 70% notes / 30% recall mixture. This dataset has
2,997,925 total loss-bearing training-data tokens
(2,098,505 notes + 899,420 recall).
The token count describes the dataset before its two training exposures.
Run:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/wmrl-v4-scale70-3m-agentic-eval-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 7.0000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think.wmrl-v4-scale70-30m-agentic-eval-20t-think
scale70-30m Agentic evaluation, 20 turns, thinking enabled
1,000 saved evaluation attempts: 250 C&H firm-knowledge tasks × four samples.
The base model is Qwen/Qwen3.5-9B at revision c202236235762e1c871ad0ccb60c8ee5ba337b9a, trained for two
epochs on a 70% notes / 30% recall mixture. This dataset has
30,004,389 total loss-bearing training-data tokens
(21,003,132 notes + 9,001,257 recall).
The token count describes the dataset before its two training exposures.
Run:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/wmrl-v4-scale70-30m-agentic-eval-20t-think.wmrl-v4-scale70-1m-agentic-eval-20t-think
scale70-1m Agentic evaluation, 20 turns, thinking enabled
1,000 saved evaluation attempts: 250 C&H firm-knowledge tasks × four samples.
The base model is Qwen/Qwen3.5-9B at revision c202236235762e1c871ad0ccb60c8ee5ba337b9a, trained for two
epochs on a 70% notes / 30% recall mixture. This dataset has
1,002,441 total loss-bearing training-data tokens
(701,671 notes + 300,770 recall).
The token count describes the dataset before its two training exposures.
Run:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/wmrl-v4-scale70-1m-agentic-eval-20t-think.
