CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ZomiLearner /English-Zomi-OPUS_Tatoeba_v20230412 English–Zomi Parallel Corpus (1.78M) This dataset contains 1.78 million English–Zomi sentence pairs, created to support machine translation, linguistic research, and large‑scale language model training. It is fully open and permissively licensed for commercial and non‑commercial use. 🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO 639‑3 code.… See the full description on the dataset page: https://huggingface.co/datasets/ZomiLearner/English-Zomi-OPUS_Tatoeba_v20230412.tabulartranslation1M<n<10M0 likes15k downloads7mo agoHugging Face02Gryphe /Opus-WritingPrompts Opus Writing Prompts This is a dataset containing 3008 short stories, generated by an unrestrained Claude Opus using Reddit's Writing Prompts as a source. Each sample is generally between 4000-6000 characters long. These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres. Disclaimer: This dataset is extremely varied and includes erotica. You have been warned. Three files are included: A ShareGPT dataset, ready to be used for… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/Opus-WritingPrompts.texttext-generation1K<n<10K86 likes7.3k downloads2y agoHugging Face03razzant /ouroboros-osworld-verified-opus5 Ouroboros on OSWorld-Verified: 90.69%, the highest result reported to date Status: Self-reported result over all 361 tasks. The official per-task scores, prompts, manifests and feasibility records are public here, together with every acting task record that the run produced. Start here Result 90.69% (327.39 / 361) Model anthropic/claude-opus-5 Method Screenshot only, one rollout, 100 policy turns Exact evidence f52ebf2 and evidence.json… See the full description on the dataset page: https://huggingface.co/datasets/razzant/ouroboros-osworld-verified-opus5.textothern<1K1 likes2.6k downloads1mo agoHugging Face04angrygiraffe /claude-opus-4.6-4.7-reasoning-8.7k Background Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed. Clarification on Reasoning The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k.texttext-generation10K<n<100K448 likes1k downloads5mo agoHugging Face05anthracite-org /kalo-opus-instruct-22k-no-refusaltext10K<n<100K38 likes717 downloads2y agoHugging Face06nothingiisreal /Claude-3-Opus-Instruct-15K Original Character Card Processed 15K Prompts - See Usable Responses Below Based on Claude 3 Opus through AWS. I took a random 5K + 10K prompt subset from Norquinal/claude_multi_instruct_30k to use as prompts, and called API for my answers. Warning! Uncleaned - Only Filtered for Blatant Refusals. I will be going through and re-prompting missing prompts, but I do not expect much success, as some of the prompts shown are nonsensical, incomplete, or impossible… See the full description on the dataset page: https://huggingface.co/datasets/nothingiisreal/Claude-3-Opus-Instruct-15K.text10K<n<100K20 likes579 downloads2y agoHugging Face07Nopm /Opus_WritingStruct Opus Writing Instruct 6k Synthetically generated creative writing data using Claude 3 Opus, by Anthropic, filtered and cleaned using automated means. Focus was placed on having as many genres as possible represented in the data, and to have Claude more openly use its excellent prose. It also contains question-answer instruction pairs related to the topic of writing. Dataset Details Curated by: Nopm License: Apache 2 Credits: The entire SillyTilly community for providing… See the full description on the dataset page: https://huggingface.co/datasets/Nopm/Opus_WritingStruct.texttext-generation1K<n<10K39 likes554 downloads2y agoHugging Face08sammshen /wildclaw-opus-traces WildClaw Agent Traces — Claude Opus 4.6 Full agentic traces from running WildClawBench tasks through Claude Opus 4.6 via an instrumented reverse proxy. Dataset Description Each trace file captures the complete HTTP-level request/response pairs between the OpenClaw agent and Claude Opus 4.6, including: System prompts, user messages, and assistant responses Tool calls and tool results (multi-turn agentic loops) Token usage and cost metadata from OpenRouter Key… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/wildclaw-opus-traces.tabularn<1K4 likes488 downloads6mo agoHugging Face09Jackrong /Claude-opus-4.7-TraceInversion-5000x 🌀 Claude-opus-4.7-TraceInversion-5000x v1.0 Release A High-Fidelity Reconstructed CoT Dataset Saturated with the 'Opus Deep Logic Style' via Trace Inversion 📊 5,000 Samples 🧬 Trace Inversion & Negentropy 🛠 SFT & DPO Ready 🔥 Claude 4.7-Max Distillation 🌐 English & Multilingual 💡 What is Trace Inversion? In Large Language Model (LLM) reasoning distillation, proprietary API models (such as GPT-4/5 and Claude)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Claude-opus-4.7-TraceInversion-5000x.texttext-generation1K<n<10K83 likes474 downloads4mo agoHugging Face10TeichAI /claude-4.5-opus-high-reasoning-250xThis is a reasoning dataset created using Claude Opus 4.5 with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated. The dataset is meant for creating distilled versions of Claude Opus 4.5 by fine-tuning already existing open-source LLMs. Stats Costs: $ 52.3 (USD) Total tokens (input + output): 2.13 M textn<1K404 likes471 downloads10mo agoHugging Face11programbench /20260731_mini-v2.4.2_opus-5-xhightextn<1K0 likes436 downloads2mo agoHugging Face12AMAImedia /NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text1M<n<10M2 likes419 downloads5d agoHugging Face13nohurry /Opus-4.6-Reasoning-3000x-filtered [!WARNING] NOTICE: The original dataset has been updated with better filtering. Please use the original dataset, not this one. Filtered from: https://huggingface.co/datasets/crownelius/Opus-4.6-Reasoning-3000x The original dataset has 979 refusals, I removed these in this version. text1K<n<10K648 likes407 downloads6mo agoHugging Face14TeichAI /lordx64-claude-opus-4.7-max-cleaned reasoning-distill-claude-opus-4-7-max-cleaned Cleaned version of lordx64/reasoning-distill-claude-opus-4-7-max. See the original dataset for full provenance, collection methodology, and terms of use. Cleaning steps Step Filter Reason Rows removed 1 Simulated thinking (...) Rows with ... in thinking/response indicate the model learned to simulate reasoning (e.g., "Now I'm laying out the puzzle grids...") rather than actually performing it. This causes failures… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/lordx64-claude-opus-4.7-max-cleaned.text1K<n<10K24 likes397 downloads5mo agoHugging Face15kalomaze /Opus_Instruct_25kA scaled-up continuation of what I did with Opus Instruct 3k. Quality checks were not very thorough so far, I did some general keyword filtering and parsing to make sure that there wasn't turn leakage. This time I made sure most responses were two turns. text10K<n<100K38 likes364 downloads2y agoHugging Face16AMAImedia /NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text10K<n<100K7 likes357 downloads5d agoHugging Face17Roman1111111 /opus-gpt-swe-frontier-core SWE Base Repository-level software engineering trajectories for training coding agents. 2,459 chat trajectories · 48,499 API calls · $837.57 recorded generation cost SWE-bench · debugging · patching · tools · agents Overview SWE Base is a software-engineering dataset centered on real repository issues. Each training example gives an agent a problem statement and captures the multi-turn process of inspecting a codebase, reasoning about a bug… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/opus-gpt-swe-frontier-core.tabulartext-generation1K<n<10K3 likes328 downloads1mo agoHugging Face18programbench /20260507_mini-v2.2.6_opus-4-7-xhightextn<1K0 likes325 downloads2mo agoHugging Face1911-47 /claude_opus_4.8_max_thinking_5k_v2 Claude Opus 4.8 MAX THINKING — Distillation Dataset 5,000 high-quality examples designed to distill the maximum-effort reasoning, honest analysis, production software engineering, and agentic capabilities of Claude Opus 4.8. Overview This dataset captures Opus 4.8’s signature strengths: Deep, structured, high-effort reasoning Honest communication about trade-offs and uncertainties Excellent production software engineering judgment Strong agentic workflow design… See the full description on the dataset page: https://huggingface.co/datasets/11-47/claude_opus_4.8_max_thinking_5k_v2.text1K<n<10K7 likes261 downloads4mo agoHugging Face20Jackrong /Claude-opus-4.6-TraceInversion-9000x 🌀 Claude-opus-4.6-TraceInversion-9000x v1.0 Release A High-Fidelity Reconstructed CoT Dataset via Trace Inversion 📊 9,000 Samples 🧬 Trace Inversion & Negentropy 🛠 SFT & DPO Ready 🔥 Claude 4.6 Distillation 🌐 English & Multilingual 💡 What is Trace Inversion? In Large Language Model (LLM) reasoning distillation, proprietary API models (such as GPT-4/5 and Claude) typically hide their internal thinking steps, providing… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Claude-opus-4.6-TraceInversion-9000x.texttext-generation1K<n<10K85 likes245 downloads4mo agoHugging Face21Roman1111111 /claude-opus-4.6-10000xThis is a high-fidelity reasoning dataset synthesized using Claude Opus 4.6. The dataset is designed to capture the model's internal "Chain of Thought" and reasoning traces, specifically focusing on mathematical accuracy and structured logical deduction. The dataset is intended for Supervised Fine-Tuning (SFT) and Distillation, allowing smaller open-source models to inherit the sophisticated reasoning patterns of Claude Opus 4.6. Dataset Description This collection combines high-difficulty… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/claude-opus-4.6-10000x.text1K<n<10K393 likes236 downloads6mo agoHugging Face22syntaxsynth /swe-bench-opus-logs Claude 3 inference SWE-Bench results Contains prompting responses from SWE-bench on these 2 settings: Oracle retrieval BM25 retrieval Each of the subsets contains an additional log_last_line attributes which is the last line from log files generated during evaluation step. Results: Model BM25 Retrieval Resolved (%) Oracle Retrieval Resolved (%) GPT-4* 0 1.74 Claude-2 1.96 4.80 Claude-3 Opus (20240229) 3.24 6.42 Claude-2 and GPT-4 results from SWE-bench… See the full description on the dataset page: https://huggingface.co/datasets/syntaxsynth/swe-bench-opus-logs.textquestion-answering1K<n<10K1 likes221 downloads3y agoHugging Face23Quaxicron /claude-opus-4.8-pi-tracesMore expensive than anticpated so you only get 4 lol :P This dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Claude Opus 4.8 Pi Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by anthropic/claude-opus-4.8. JSONL files: 4 Training-ready tools A complete configured tools schema snapshot is… See the full description on the dataset page: https://huggingface.co/datasets/Quaxicron/claude-opus-4.8-pi-traces.tabulartext-generationn<1K0 likes221 downloads3mo agoHugging Face24armand0e /badlogicgames-pi-mono-opus-filteredFiltered version of badlogicgames/pi-mono - Only opus traces, dropped invalid sessions as well. All traces present are training safe and teich compatible tabularn<1K2 likes197 downloads4mo agoHugging Face25programbench /20260429_mini-v2.2.6_opus-4-7textn<1K0 likes190 downloads3mo agoHugging Face26PocketDoc /Dans-Prosemaxx-Opus-Writingtextn<1K1 likes189 downloads2y agoHugging Face27armand0e /claude-opus-4.8-pi-tracesMore expensive than anticpated so you only get 4 lol :P This dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Claude Opus 4.8 Pi Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by anthropic/claude-opus-4.8. JSONL files: 4 Training-ready tools A complete configured tools schema snapshot is… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/claude-opus-4.8-pi-traces.tabulartext-generationn<1K8 likes143 downloads4mo agoHugging Face28kalomaze /Opus_Instruct_3k~2.5 million tokens worth of general purpose multi-turn instruction finetuning data in the style of Claude 3 Opus. Turns for both the Human and Assistant are created by the model itself (ala Magpie). text1K<n<10K27 likes141 downloads2y agoHugging Face29emanubiz /opus-agent-merged-v2text1K<n<10K0 likes138 downloads5mo agoHugging Face30Mahfug /claude-opus-4.6-4.7-reasoning-8.7k Background Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed. Clarification on Reasoning The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking" expected to… See the full description on the dataset page: https://huggingface.co/datasets/Mahfug/claude-opus-4.6-4.7-reasoning-8.7k.texttext-generation10K<n<100K3 likes128 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.