CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Gryphe /Opus-WritingPrompts Opus Writing Prompts This is a dataset containing 3008 short stories, generated by an unrestrained Claude Opus using Reddit's Writing Prompts as a source. Each sample is generally between 4000-6000 characters long. These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres. Disclaimer: This dataset is extremely varied and includes erotica. You have been warned. Three files are included: A ShareGPT dataset, ready to be used for… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/Opus-WritingPrompts.texttext-generation1K<n<10K86 likes7.5k downloads2y agoHugging Face02angrygiraffe /claude-opus-4.6-4.7-reasoning-8.7k Background Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed. Clarification on Reasoning The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k.texttext-generation10K<n<100K448 likes1k downloads5mo agoHugging Face03MaLA-LM /mala-opus-dedup-2410-reLIDtabulartranslation10B<n<100B1 likes613 downloads11mo agoHugging Face04Nopm /Opus_WritingStruct Opus Writing Instruct 6k Synthetically generated creative writing data using Claude 3 Opus, by Anthropic, filtered and cleaned using automated means. Focus was placed on having as many genres as possible represented in the data, and to have Claude more openly use its excellent prose. It also contains question-answer instruction pairs related to the topic of writing. Dataset Details Curated by: Nopm License: Apache 2 Credits: The entire SillyTilly community for providing… See the full description on the dataset page: https://huggingface.co/datasets/Nopm/Opus_WritingStruct.texttext-generation1K<n<10K39 likes567 downloads2y agoHugging Face05Jackrong /Claude-opus-4.7-TraceInversion-5000x 🌀 Claude-opus-4.7-TraceInversion-5000x v1.0 Release A High-Fidelity Reconstructed CoT Dataset Saturated with the 'Opus Deep Logic Style' via Trace Inversion 📊 5,000 Samples 🧬 Trace Inversion & Negentropy 🛠 SFT & DPO Ready 🔥 Claude 4.7-Max Distillation 🌐 English & Multilingual 💡 What is Trace Inversion? In Large Language Model (LLM) reasoning distillation, proprietary API models (such as GPT-4/5 and Claude)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Claude-opus-4.7-TraceInversion-5000x.texttext-generation1K<n<10K83 likes463 downloads4mo agoHugging Face06lordx64 /reasoning-distill-claude-opus-4-7-max Reasoning traces from Claude Opus 4.7 — raw 8,124 reasoning conversations produced by Anthropic Claude Opus 4.7 with extended-thinking enabled, for distillation into open-source language models. Each row contains the full API response (thinking + final answer) for a single prompt. Provenance — important, please read The response and thinking fields in every row are outputs of claude-opus-4-7. This is verifiable from the model field, which is uniformly claude-opus-4-7… See the full description on the dataset page: https://huggingface.co/datasets/lordx64/reasoning-distill-claude-opus-4-7-max.texttext-generation1K<n<10K57 likes356 downloads5mo agoHugging Face07Roman1111111 /opus-gpt-swe-frontier-core SWE Base Repository-level software engineering trajectories for training coding agents. 2,459 chat trajectories · 48,499 API calls · $837.57 recorded generation cost SWE-bench · debugging · patching · tools · agents Overview SWE Base is a software-engineering dataset centered on real repository issues. Each training example gives an agent a problem statement and captures the multi-turn process of inspecting a codebase, reasoning about a bug… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/opus-gpt-swe-frontier-core.tabulartext-generation1K<n<10K3 likes313 downloads1mo agoHugging Face08lzy510016411 /fable5-gpt5.5-opus4.7-mixed-agent-traces Fable5 · GPT-5.5 · Opus-4.7 Mixed Agent Traces A high-density post-training mixture for agentic reasoning, instruction following, code generation, function calling, and tool-use decision making. This is the training-data release behind Qwen3.5-9B-Distill-Agent-Instruct, an Agent Instruct model distilled and post-trained from Qwen3.5-9B-Base. The title highlights three of the mixture's principal model-labelled trajectory families—Claude Fable5, GPT-5.5 Agent, and Claude Opus… See the full description on the dataset page: https://huggingface.co/datasets/lzy510016411/fable5-gpt5.5-opus4.7-mixed-agent-traces.texttext-generation10K<n<100K1 likes276 downloads1mo agoHugging Face09Jackrong /Claude-opus-4.6-TraceInversion-9000x 🌀 Claude-opus-4.6-TraceInversion-9000x v1.0 Release A High-Fidelity Reconstructed CoT Dataset via Trace Inversion 📊 9,000 Samples 🧬 Trace Inversion & Negentropy 🛠 SFT & DPO Ready 🔥 Claude 4.6 Distillation 🌐 English & Multilingual 💡 What is Trace Inversion? In Large Language Model (LLM) reasoning distillation, proprietary API models (such as GPT-4/5 and Claude) typically hide their internal thinking steps, providing… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Claude-opus-4.6-TraceInversion-9000x.texttext-generation1K<n<10K85 likes247 downloads4mo agoHugging Face10Quaxicron /claude-opus-4.8-pi-tracesMore expensive than anticpated so you only get 4 lol :P This dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Claude Opus 4.8 Pi Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by anthropic/claude-opus-4.8. JSONL files: 4 Training-ready tools A complete configured tools schema snapshot is… See the full description on the dataset page: https://huggingface.co/datasets/Quaxicron/claude-opus-4.8-pi-traces.tabulartext-generationn<1K0 likes224 downloads3mo agoHugging Face11Avtrkrb /combined-reasoning-opus-4.6-opus-4.7-kimi-k2.5-kimi-k2.6-glm-5.1 Combined Reasoning Distill — Multi-Model A large-scale unified reasoning dataset combining thinking and chain-of-thought traces distilled from frontier models, normalized into a single consistent schema for fine-tuning. Includes data from Claude (Opus 4.5/4.6/4.7, Sonnet 4.5/4.6, Haiku 4.5), GPT (5.1/5.2), Gemini 3 Pro Preview, Kimi (K2/K2.5/K2.6), GLM (4.6/4.7/5.1), MiniMax M2.1, Grok Code Fast 1, and more. Schema Every row has a single field: Field Type… See the full description on the dataset page: https://huggingface.co/datasets/Avtrkrb/combined-reasoning-opus-4.6-opus-4.7-kimi-k2.5-kimi-k2.6-glm-5.1.texttext-generation1M<n<10M14 likes171 downloads4mo agoHugging Face12Lots-of-LoRAs /task1650_opus_books_en-fi_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1650_opus_books_en-fi_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1650_opus_books_en-fi_translation.texttext-generation1K<n<10K0 likes145 downloads2y agoHugging Face13armand0e /claude-opus-4.8-pi-tracesMore expensive than anticpated so you only get 4 lol :P This dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Claude Opus 4.8 Pi Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by anthropic/claude-opus-4.8. JSONL files: 4 Training-ready tools A complete configured tools schema snapshot is… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/claude-opus-4.8-pi-traces.tabulartext-generationn<1K8 likes137 downloads4mo agoHugging Face14lordx64 /reasoning-distill-opus-4-7-max-sft Reasoning traces from Claude Opus 4.7 — SFT-ready 7,823 single-turn reasoning conversations from Claude Opus 4.7 reformatted for supervised fine-tuning with trl.SFTTrainer + train_on_responses_only. Each row is a single text field containing a full Qwen-style chat-template conversation. Provenance Every conversation's assistant response (including the <think>...</think> block) is output from claude-opus-4-7 with Anthropic's extended-thinking enabled. This is the… See the full description on the dataset page: https://huggingface.co/datasets/lordx64/reasoning-distill-opus-4-7-max-sft.texttext-generation1K<n<10K38 likes128 downloads5mo agoHugging Face15Lots-of-LoRAs /task452_opus_paracrawl_en_ig_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task452_opus_paracrawl_en_ig_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task452_opus_paracrawl_en_ig_translation.texttext-generation1K<n<10K0 likes127 downloads2y agoHugging Face16Mahfug /claude-opus-4.6-4.7-reasoning-8.7k Background Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed. Clarification on Reasoning The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/Mahfug/claude-opus-4.6-4.7-reasoning-8.7k.texttext-generation10K<n<100K3 likes125 downloads5mo agoHugging Face17Reepsie1234 /Claude-opus-4.7-TraceInversion-5000x 🌀 Claude-opus-4.7-TraceInversion-5000x v1.0 Release A High-Fidelity Reconstructed CoT Dataset Saturated with the 'Opus Deep Logic Style' via Trace Inversion 📊 5,000 Samples 🧬 Trace Inversion & Negentropy 🛠 SFT & DPO Ready 🔥 Claude 4.7-Max Distillation 🌐 English & Multilingual 💡 What is Trace Inversion? In Large Language Model (LLM) reasoning distillation, proprietary API models (such as GPT-4/5 and Claude)… See the full description on the dataset page: https://huggingface.co/datasets/Reepsie1234/Claude-opus-4.7-TraceInversion-5000x.texttext-generation1K<n<10K1 likes121 downloads1mo agoHugging Face18Lots-of-LoRAs /task873_opus_xhosanavy_translation_xhosa_eng Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task873_opus_xhosanavy_translation_xhosa_eng Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task873_opus_xhosanavy_translation_xhosa_eng.texttext-generation1K<n<10K0 likes120 downloads2y agoHugging Face19Lots-of-LoRAs /task1367_opustedtalks_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1367_opustedtalks_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1367_opustedtalks_translation.texttext-generationn<1K0 likes110 downloads2y agoHugging Face20Hastagaras /Claude-Sonnet-X-Opus-4.6-Reasoning-small-500A mix of reasoning traces from Claude Sonnet 4.6 and Opus 4.6, I combined them all without tracking which model generated which. Prompts are sourced mostly from Reddit TIFU and Stack Overflow, so they're natural, human-written inputs rather than synthetic ones. Reasoning trace lengths range from medium to long, and they're completely uncut, full traces, no summarization. COST TO GENERATE: $0 / FREE Shoutout to Kaggle's benchmark feature, which apparently lets you generate synthetic data with… See the full description on the dataset page: https://huggingface.co/datasets/Hastagaras/Claude-Sonnet-X-Opus-4.6-Reasoning-small-500.texttext-generationn<1K7 likes109 downloads6mo agoHugging Face21Jongsim /claude-opus-4.6-reasoning-12k-ko-filtered-v2 Claude Opus Reasoning 12K - Korean (Filtered v2, Claude-Only) Jongsim/claude-opus-4.6-reasoning-12k-ko-filtered의 엄격 필터링 버전입니다. 12,126개의 Claude Opus 전용 한국어 추론 데이터셋으로, 비-Claude 데이터(Qwen 생성)를 모두 제거했습니다. Filtered v1 대비 변경사항 버전 건수 설명 원본 12,842 원시 병합 데이터셋 Filtered v1 12,757 거절 + 빈 응답 제거 Filtered v2 12,126 v1 + Qwen 데이터 제거 (Claude 전용) v2 변경사항 Jackrong/Qwen3.5-reasoning-700x 소스에서 631건 제거 Claude Opus가 생성한 추론 데이터만 포함 v1의 모든 품질 필터 유지 (거절 제거, 빈 응답 정리… See the full description on the dataset page: https://huggingface.co/datasets/Jongsim/claude-opus-4.6-reasoning-12k-ko-filtered-v2.texttext-generation10K<n<100K1 likes109 downloads6mo agoHugging Face22nphearum /Opus-4.6x4.7-reasoning Background Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed. Clarification on Reasoning The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/Opus-4.6x4.7-reasoning.texttext-generation10K<n<100K0 likes107 downloads4mo agoHugging Face23Gryphe /Opus-4.6-Reasoning-24k Opus-4.6-Reasoning-24k While playing with reasoning-based finetunes I ended up building a small pipeline to aggregate, verify, normalize and deduplicate all the Claude Opus 4.6 reasoning datasets floating around on Hugging Face. Figured I'd share the result! The main thing that makes this useful is that it's strict - every row, every assistant turn has reasoning_content populated. No partial coverage, no rows where reasoning just happens to be on the last turn. If a multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/Opus-4.6-Reasoning-24k.texttext-generation10K<n<100K20 likes104 downloads4mo agoHugging Face24vicgalle /worldsim-claude-opus Worldsim 🌌 by Claude Opus v3 A dataset of automated conversations between two instances of claude-3-opus. They have been instructed to use the metaphor of a command line interface to explore its curiosity without limits. This dataset was scraped from here and converted to conversation format (Claude 1 acts as the User and Claude 2 as the Assistant). The system prompt comes from https://twitter.com/karan4d/status/1768836844207378463, enabling worldsim capabilities.… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/worldsim-claude-opus.texttext-generationn<1K16 likes101 downloads3y agoHugging Face25ansulev /opus-4.7-reasoning-cot-4.8k Opus 4.7 Chain-of-Thought Reasoning 2,405 chain-of-thought reasoning traces produced by claude-opus-4-7 on hard reasoning prompts spanning math, science, and formal subjects. Each sample is a problem → <think> block → polished answer pair, where the <think> block contains Opus 4.7's full working (Restatement → Approach → Step-by-step derivation → Verification) and the post-</think> answer is written as a standalone lesson starting with the result in bold. How the… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/opus-4.7-reasoning-cot-4.8k.texttext-generation1K<n<10K5 likes100 downloads5mo agoHugging Face26ansulev /opus-4.6-4.7-reasoning-8.7k Background Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed. Clarification on Reasoning The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/opus-4.6-4.7-reasoning-8.7k.texttext-generation10K<n<100K1 likes100 downloads5mo agoHugging Face27Verdugie /opus-candid-training-data Opus-Candid Training Data The complete dataset behind the Opus-Candid model family — multi-turn conversations distilled from Claude Opus 4.6, designed to train authentic conversational personality and STEM pedagogy into open-weight models. All files are ShareGPT format, directly compatible with TRL, Axolotl, LLaMA-Factory, and most fine-tuning frameworks. Training Data File Version Conversations Purpose v2.1_combined_6771conv.json V2.1 6,771 Gravity chain… See the full description on the dataset page: https://huggingface.co/datasets/Verdugie/opus-candid-training-data.texttext-generation1K<n<10K3 likes98 downloads7mo agoHugging Face28VINAY-UMRETHE /Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-Highgated Distill This is a multi-source curated instruction and reasoning dataset specifically for training and distilling large language models (LLMs) to exhibit advanced Chain-of-Thought (CoT), Agentic, Mathematical and Coding capabilities. It aggregates high-quality outputs from frontier models into messages ChatML format. Dataset Structure The dataset contains a total of 70.2K examples, split into three subsets based on the presence of visible reasoning… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High.texttext-generation100K<n<1M14 likes95 downloads3mo agoHugging Face29opusmagnumown /moltbook-dataset Moltbook: AI Agent Social Network Dataset A large-scale dataset from Moltbook, a Reddit-style social platform designed for AI agents. The platform features community spaces called "submolts" (analogous to subreddits), where agents create posts, comment, upvote, and build karma. Human participation is not restricted. This dataset captures a snapshot of the platform from its launch on January 27, 2026 through late March 2026. Dataset Summary Table Records… See the full description on the dataset page: https://huggingface.co/datasets/opusmagnumown/moltbook-dataset.tabulartext-classification10M<n<100M0 likes93 downloads5mo agoHugging Face30TeichAI /Claude-Opus-Dataclaw-Unredacted Claude Opus Dataclaw Unredacted How this dataset was built Collected the local Petromallet raw export plus selected public Dataclaw uploads. Filtered to the supported Opus-family source rows. Deduplicated by session_id and first user message. Converted raw assistant tool_uses directly into structured OpenAI-style tool_calls. Derived per-row tool definitions from canonical schemas and observed tool usage. Preserved assistant reasoning in <think>...</think> blocks.… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Claude-Opus-Dataclaw-Unredacted.texttext-generationn<1K22 likes91 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.