CoolFace
Datasetpublic

Rooftech650/claude-opus-4.6-4.7-reasoning-8.7k

Background Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed. Clarification on Reasoning The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/Rooftech650/claude-opus-4.6-4.7-reasoning-8.7k.

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes60downloads
Dataset Card

Background

Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.

Clarification on Reasoning

The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to reach the Assistant response.

There are now non-reasoning versions as well.

Files

Four datasets provided: | Split | File | Examples | Contents | |-------|------|---------:|----------| | Full | full_train.jsonl | 8,706 | All examples across all 28 categories. | | Instruct | instruct_train.jsonl | 7,217 | All 24 instructional categories — coding, math, sciences, humanities, arts, finance, medicine, law, business, linguistics, creative writing, general. | | Roleplay | roleplay_train.jsonl | 1,489 | The four creative categories — roleplay_hero, roleplay_villain, roleplay_crossover, narrative_prose. | | Code | code_train.jsonl | 1,840 | coding + math only. For coding/math-focused fine-tunes. |

SLOP Readme

A synthetic instruction-tuning dataset designed to teach language models how to think, not just what to say. Every example is written to expert standards across 28 categories spanning coding, math, the sciences, the humanities, the arts, finance, medicine, law, business, linguistics, creative writing, roleplay, and narrative prose. Every assistant turn includes a <think> block — genuine deliberation, not a reformatted answer.

Dataset Summary

MetricValue
Total examples8,706
Estimated tokens~17.0M
Categories28 (all populated)
With reasoning tags8,706 (100%)
Multi-turn conversations3,454 (39.7%)
Unique system prompts5,814
FormatOpenAI chat (JSONL)
LanguageEnglish
Teacher modelsClaude Opus 4.6 (53.7%), Claude Opus 4.7 (46.3%)

What Makes This Dataset Different

  • —Genuine chain-of-thought reasoning, on every example. Each assistant turn opens with a <think>...</think> block of 150–500 words showing the model considering multiple angles, weighing alternatives, and planning response structure — not "Step 1, Step 2" reformulations of the answer.
  • —Expert-level depth. Responses are written to satisfy senior practitioners. Coding answers explain design trade-offs. History answers engage with historiographical debate. Creative-writing critique includes line-level rewrites. Roleplay characters have internally coherent worldviews.
  • —Natural user voice. User messages sound like real people — frustrated developers pasting broken code, students challenging an explanation, novelists stuck mid-draft, editors asking for a tonal shift. Hard rule: at most ~20% of user messages start with What or How.
  • —5,814 unique system prompts. Domain-specific personas (e.g. "You are a database performance consultant working on a Postgres query that's hitting timeouts under load") rather than one generic "helpful assistant" repeated thousands of times.
  • —Character-accurate roleplay. Roleplay examples are built around source-material voice, verbal habits, and worldview — not surface costumes. Includes a deliberately dark track of villain, hero, and crossover examples written in the literary register of Le Carré, McCarthy, Atwood, Tartt, Flynn, Bakker, and similar reference points.
  • —No refusals or safety hedging. Refusals, content warnings, and clarification-only turns are intentionally excluded. This dataset is for teaching capability, not for replacing alignment training.

Categories

28 categories grouped into instructional and creative/roleplay sets. All 28 are populated; the largest categories carry the foundational legacy content, while the newer per-discipline categories give per-domain coverage.

Instructional categories (24)

CategoryDescription
codingWorking code with design trade-offs, debugging, architecture. Python, TypeScript, Go, Rust, SQL, more.
mathPure and applied mathematics, statistics, probability, geometry, algebra, calculus, logic.
physicsMechanics, thermodynamics, quantum, relativity, electromagnetism, optics.
biologyGenetics, evolution, ecology, microbiology, neuroscience, cell biology.
chemistryOrganic, inorganic, biochemistry, materials science.
earth_scienceGeology, climate, meteorology, oceanography, astronomy, paleontology.
scienceGeneral-science catch-all for cross-disciplinary topics.
historyEvents, historiography, primary sources, ancient through modern.
philosophyEpistemology, ethics, logic, metaphysics, aesthetics.
psychologyCognition, behavior, development, social psychology.
political_scienceGovernance, international relations, policy, political theory.
sociologySocial structures, institutions, inequality, demography.
economicsMacro/microeconomics, econometrics, development economics, game theory.
geographyHuman and physical geography, cartography, geopolitics, urban planning.
literatureLiterary criticism, poetry analysis, comparative literature, theory.
humanitiesCatch-all for cross-disciplinary humanities topics.
artsMusic, film, theater, painting, sculpture, architecture, photography, design.
financeInvesting, accounting, banking, markets, personal finance, trading.
medicineClinical reasoning, pharmacology, anatomy, public health, epidemiology.
lawConstitutional, contracts, criminal, civil, jurisprudence, regulation.
businessManagement, strategy, leadership, operations, marketing, entrepreneurship.
linguisticsTranslation, etymology, phonetics, grammar, syntax, language acquisition.
creative_writingCraft-focused coaching with concrete techniques, before/after rewrites, line-level analysis.
generalPractical advice, explanations, life questions. Depth matched to question complexity.

Creative / roleplay categories (4)

CategoryDescription
roleplay_heroHeroic and morally complex protagonists with rich, source-accurate voices.
roleplay_villainAntagonists with internally coherent worldviews — not cartoonish evil.
roleplay_crossoverCross-canon character pairings with distinct voices and dramatic dynamics.
narrative_prosePublishable-quality literary fiction in named author voices (Hemingway, Tolstoy, Austen, Pynchon, McCarthy, Le Carré, etc.) and genres.

Overall

MetricValue
Examples8,706
Tokens (estimated)17,013,533
Avg tokens / example1,954
With reasoning8,706 (100.0%)
Multi-turn3,454 (39.7%)
Single-turn5,252 (60.3%)

Category Counts

CategoryExamplesTokensMulti-turn %
coding1,6282,545,22130.4%
humanities8621,849,70832.5%
science7371,681,34637.4%
roleplay_hero419640,08463.5%
roleplay_villain378635,98460.8%
narrative_prose377710,80743.0%
roleplay_crossover315581,18856.8%
creative_writing281532,50430.6%
medicine280519,66222.1%
biology277541,01321.3%
general276284,69637.0%
arts245576,17041.2%
chemistry221508,54652.9%
physics220512,19656.8%
math212394,90754.2%
geography155358,32142.6%
history155348,82241.3%
economics155380,37242.6%
political_science154374,90138.3%
sociology154378,26142.2%
business152315,06538.2%
earth_science152358,20941.4%
finance151328,60738.4%
philosophy150335,51441.3%
linguistics150306,88939.3%
literature150299,60638.7%
psychology150339,56539.3%
law150375,36041.3%

Per-category JSONL splits live in categories/.

By Model

Every example carries a model field identifying which Claude model generated it.

ModelCountShareTokens
claude-opus-4-64,67553.7%6,304,169
claude-opus-4-74,03146.3%10,709,363

The two model populations are roughly balanced by example count, but Opus 4.7 examples carry ~70% more tokens on average — newer waves trend toward longer multi-turn content.

Turn Distribution

TurnsExamples%
15,25260.3%
21,49117.1%
31,85821.3%
4820.9%
5210.2%
620.0%

Multi-turn conversations are designed to teach models to build on context, handle follow-ups that change direction, defend a craft choice, revise on request, and adjust depth based on user response.

Response Length Distribution

Assistant message length, in characters:

PercentileCharacters
p102,061
p252,914
Median4,239
p755,682
p907,052
Max30,026

Reasoning blocks themselves are typically 150–500 words; the rest is the user-facing answer.

Format

Standard OpenAI chat format in JSONL. Each line is one JSON object with category, messages, and model fields:

json
{
  "category": "coding",
  "model": "claude-opus-4-7",
  "messages": [
    {"role": "system", "content": "You are a senior backend engineer reviewing performance issues..."},
    {"role": "user", "content": "Explain move semantics to me..."},
    {"role": "assistant", "content": "<reasoning>\nThe user understands C++ fundamentals but...\n</reasoning>\n\n`std::move` does not move anything. It is a cast..."}
  ]
}

The category and model fields are metadata for filtering and provenance — fine-tuning APIs read only messages.

Terms

Use it for things you should use it for but don't use it for anything you shouldn't use it for. Like Anthropic does, always respect terms of use...