datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k.Opus_WritingStruct
Opus Writing Instruct 6k
Synthetically generated creative writing data using Claude 3 Opus, by Anthropic, filtered and cleaned using automated means. Focus was placed on having as many genres as possible represented in the data, and to have Claude more openly use its excellent prose.
It also contains question-answer instruction pairs related to the topic of writing.
Dataset Details
Curated by: Nopm
License: Apache 2
Credits: The entire SillyTilly community for providing… See the full description on the dataset page: https://huggingface.co/datasets/Nopm/Opus_WritingStruct.swe-bench-opus-logs
Claude 3 inference SWE-Bench results
Contains prompting responses from SWE-bench on these 2 settings:
Oracle retrieval
BM25 retrieval
Each of the subsets contains an additional log_last_line attributes which is the last line from log files generated during evaluation step.
Results:
Model
BM25 Retrieval Resolved (%)
Oracle Retrieval Resolved (%)
GPT-4*
0
1.74
Claude-2
1.96
4.80
Claude-3 Opus (20240229)
3.24
6.42
Claude-2 and GPT-4 results from SWE-bench… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/swe-bench-opus-logs.opus-4.7-reasoning-cot-4.8k
Opus 4.7 Chain-of-Thought Reasoning
2,405 chain-of-thought reasoning traces produced by claude-opus-4-7 on hard reasoning prompts spanning math, science, and formal subjects.
Each sample is a problem → <think> block → polished answer pair, where the <think> block contains Opus 4.7's full working (Restatement → Approach → Step-by-step derivation → Verification) and the post-</think> answer is written as a standalone lesson starting with the result in bold.
How the… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/opus-4.7-reasoning-cot-4.8k.Opus-4.6x4.7-reasoning
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/Opus-4.6x4.7-reasoning.opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/opus-4.6-4.7-reasoning-8.7k.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/Mahfug/claude-opus-4.6-4.7-reasoning-8.7k.Opus-4.6-RU-Reasoning-creative-1385x-not-filtered
Opus-4.6-RU-Creative-Writing — Russian Creative Writing Reasoning Dataset
A Russian-language dataset of creative writing tasks generated with Claude claude-opus-4.6 (extended thinking enabled). Each sample contains a creative prompt, a full reasoning chain showing the creative process, and a detailed artistic response.
Dataset Info
Language: Russian 🇷🇺
Size: ~1,385 samples (growing)
Model used: anthropic/claude-opus-4.6 with reasoning: {effort: "high"}
Format:… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/Opus-4.6-RU-Reasoning-creative-1385x-not-filtered.Claude-Sonnet-Opus
🧠 Claude Sonnet + Opus (Gemma 4 Reasoning Dataset)
A massive, high-quality analytical reasoning dataset built from Claude Sonnet 4.6 and Claude Opus 4.6/4.7. This dataset has been forensically scrubbed of all system prompts, AI personas, and roleplay—leaving behind a pure, highly-distilled engine for teaching Gemma 4 how to think.
⚡ Why This Dataset is Different
Raw Claude datasets often contain baked-in system prompts like "You are Claude, created… See the full description on the dataset page: https://huggingface.co/datasets/qsardor/Claude-Sonnet-Opus.Opus-4.5-3000x
Claude Opus 4.5 3000x Dataset
Dataset Description
A premium dataset containing 3,096 high-quality samples generated using Claude Opus 4.5, Anthropic's most capable model. This dataset emphasizes reasoning, creative writing, and mathematical problem-solving.
Dataset Summary
Total Samples: 3,096
Model: Claude Opus 4.5
Languages: English
Format: JSON
License: Apache 2.0
Quality: Premium (from Anthropic's flagship model)
Task Distribution… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Opus-4.5-3000x.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/Mattral/claude-opus-4.6-4.7-reasoning-8.7k.opus-4.7-reasoning-cot
Opus 4.7 Chain-of-Thought Reasoning
2,405 chain-of-thought reasoning traces produced by claude-opus-4-7 on hard reasoning prompts spanning math, science, and formal subjects.
Each sample is a problem → <think> block → polished answer pair, where the <think> block contains Opus 4.7's full working (Restatement → Approach → Step-by-step derivation → Verification) and the post-</think> answer is written as a standalone lesson starting with the result in bold.
How the data… See the full description on the dataset page: https://huggingface.co/datasets/eddieran/opus-4.7-reasoning-cot.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/jotbruh2/claude-opus-4.6-4.7-reasoning-8.7k.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/felycia/claude-opus-4.6-4.7-reasoning-8.7k.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/pctoby/claude-opus-4.6-4.7-reasoning-8.7k.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/etbbebe/claude-opus-4.6-4.7-reasoning-8.7k.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/manojdahal191gom/claude-opus-4.6-4.7-reasoning-8.7k.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/thongfamilynguyen1126/claude-opus-4.6-4.7-reasoning-8.7k.claude-opus-4-8-xhigh-reasoning-8.7k
Background
This is a Claude-Opus-4.8-xhigh upgrade for angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k. Generated with Claude API.
How the data is synthesised?
Every example is produced by a two-pass "answer-first, then reason-back" pipeline against the Claude API (claude-opus-4-8), with adaptive thinking on and effort: xhigh for both passes. The reasoning you see in each <think> block is synthetic — a first-person deliberation written to plausibly lead to the… See the full description on the dataset page: https://huggingface.co/datasets/Met4physics/claude-opus-4-8-xhigh-reasoning-8.7k.Opus-4.6-Reasoning-2160x
Opus-4.6-Reasoning-2160x
2,160 high-quality reasoning traces generated by Claude Opus 4.6 via OpenRouter, covering mathematics, competitive programming, logic, science, and language tasks. Each example includes the full problem, an extended chain-of-thought, and a final solution — making the dataset suitable for supervised fine-tuning, chain-of-thought distillation, and reasoning-capability transfer to smaller models.
Originally generated as a batch of 3,305 examples; 1,145 were… See the full description on the dataset page: https://huggingface.co/datasets/Glint-Research/Opus-4.6-Reasoning-2160x.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/angin1920/claude-opus-4.6-4.7-reasoning-8.7k.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/Rooftech650/claude-opus-4.6-4.7-reasoning-8.7k.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/Hiren122/claude-opus-4.6-4.7-reasoning-8.7k.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/txchmechanicus/claude-opus-4.6-4.7-reasoning-8.7k.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/pmshal232/claude-opus-4.6-4.7-reasoning-8.7k.BaaderSo36-Opus4.7-REAP
BaaderSo36-Opus4.7-REAP
A synthetic reasoning dataset generated using Anthropic Claude Opus 4.7 (claude-opus-4-7). Each sample contains an explicit <think>...</think> reasoning block followed by a Final answer: boundary and the actual response.
Dataset Statistics
Total samples: 1739
Source distribution:
debug: 604
react_advanced: 421
math_hard: 260
humaneval: 164
code_contests: 135
math_l5: 125
react: 30
Reasoning depth (characters in thinking block)… See the full description on the dataset page: https://huggingface.co/datasets/baaderso36/BaaderSo36-Opus4.7-REAP.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/Hein1212/claude-opus-4.6-4.7-reasoning-8.7k.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/atelLex/claude-opus-4.6-4.7-reasoning-8.7k.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/hardcoremoore/claude-opus-4.6-4.7-reasoning-8.7k.Opus-4.6-reasoning-sft-12k
Opus-4.6-reasoning-sft-12k
A unified conversational reasoning dataset built by combining and normalizing two Hugging Face datasets into a single training-ready schema for supervised fine-tuning.
Overview
This dataset was created to support reasoning-focused SFT for chat models, especially Qwen-family conversational models and derivatives.
It unifies two source datasets into one consistent format:
Roman1111111/claude-opus-4.6-10000x
Crownelius/Opus-4.6-Reasoning-3300x… See the full description on the dataset page: https://huggingface.co/datasets/ykarout/Opus-4.6-reasoning-sft-12k.
