CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01greghavens /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K65 likes1.6k downloads2mo agoHugging Face02greghavens /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K21 likes705 downloads2mo agoHugging Face03DSFFGFG456 /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,380 TRAJECTORIES · 12,490 TRAINING ROWS · 14 MB PARQUET · 663 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/DSFFGFG456/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K3 likes543 downloads2mo agoHugging Face04jiajiale9 /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 697 TRAJECTORIES · 4,890 TRAINING ROWS · 3 MB PARQUET · 89 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/jiajiale9/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes277 downloads2mo agoHugging Face05ArkhAngelLifeJiggy /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes161 downloads2mo agoHugging Face06moehamid /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,374 TRAJECTORIES · 12,448 TRAINING ROWS · 14 MB PARQUET · 662 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/moehamid/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K0 likes139 downloads2mo agoHugging Face07laion /nemotron-terminal-debugging nemotron-terminal-debugging Per-source partition of nvidia/Nemotron-Terminal-Corpus, filtered to source == "debugging". The difficulty column preserves the original easy / medium / mixed split (na for the dataset_adapters/* files, which did not carry a difficulty label). Partitioning scheme: adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet {skill} (e.g. debugging, security, …) — rows from synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-debugging.textquestion-answering10K<n<100K1 likes138 downloads5mo agoHugging Face08greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes136 downloads2mo agoHugging Face09moehamid /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 601 TRAJECTORIES · 4,089 TRAINING ROWS · 3 MB PARQUET · 73 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/moehamid/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes133 downloads2mo agoHugging Face10siddharth0713 /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,161 TRAJECTORIES · 11,235 TRAINING ROWS · 11 MB PARQUET · 656 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/siddharth0713/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K0 likes114 downloads2mo agoHugging Face11Distillio /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/Distillio/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes104 downloads1mo agoHugging Face1211-47 /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/11-47/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes103 downloads5d agoHugging Face13gbeck /kimi-k3-coding-and-debugging-traces Kimi K3 Coding & Debugging Agent Traces Generated by moonshiner — an open harness for distilling verified, model-attested agentic coding traces. Real, end-to-end agentic coding trajectories produced by moonshotai/kimi-k3 driving the pi coding-agent runtime over openrouter, at max reasoning. Each trajectory solves a concrete repair or build task in a real repository — reading, editing, and running code with tools — and is published only after its work verifiably passes —… See the full description on the dataset page: https://huggingface.co/datasets/gbeck/kimi-k3-coding-and-debugging-traces.texttext-generationn<1K3 likes100 downloads2mo agoHugging Face14ansulev /gpt-5-6-sol-coding-and-debugging-traces Mirror: greghavens/gpt-5.6-sol-coding-and-debugging-traces Pinned snapshot / mirror of greghavens/gpt-5.6-sol-coding-and-debugging-traces, re-hosted for PROTISEC research reproducibility. Redistributed under the upstream license (cc-by-4.0) with attribution — all credit to the original author. Original author: greghavens Source dataset: greghavens/gpt-5.6-sol-coding-and-debugging-traces License: cc-by-4.0 Family: coding_traces Mode: stream Rows cached: 17939 Changes vs… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/gpt-5-6-sol-coding-and-debugging-traces.tabular10K<n<100K0 likes100 downloads2mo agoHugging Face15taisazero /socratic-debugging-benchmark Socratic Debugging Benchmark The repository contains the dataset for the Socratic Debugging Benchmark accompanying the papers "Socratic Questioning of Novice Debuggers: A Benchmark Dataset and Preliminary Evaluations" in proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Application at ACL 2023 and "Can Language Models Employ the Socratic Method? Experiments with Code Debugging" in the proceedings of SIGCSE'24. The dataset is also hosted on… See the full description on the dataset page: https://huggingface.co/datasets/taisazero/socratic-debugging-benchmark.texttext-generationn<1K2 likes90 downloads3y agoHugging Face16rashdan1 /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/rashdan1/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes83 downloads2mo agoHugging Face17ansulev /kimi-k3-coding-and-debugging-traces Mirror: greghavens/kimi-k3-coding-and-debugging-traces Pinned snapshot / mirror of greghavens/kimi-k3-coding-and-debugging-traces, re-hosted for PROTISEC research reproducibility. Redistributed under the upstream license (cc-by-4.0) with attribution — all credit to the original author. Original author: greghavens Source dataset: greghavens/kimi-k3-coding-and-debugging-traces License: cc-by-4.0 Family: coding_traces Mode: stream Rows cached: 4607 Changes vs upstream: cached… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/kimi-k3-coding-and-debugging-traces.tabular1K<n<10K0 likes81 downloads2mo agoHugging Face18creeperdatasets /python_debugging Python Debugging A synthetic instruction-tuning dataset for training AI models to identify and fix bugs in Python code. Dataset Summary Field Value Entries 75 Format input / output pairs Language English Topic Finding and fixing bugs in Python code Synthetic Yes, generated with DeepSeek License MIT Dataset Description Each entry presents a snippet of Python code containing a deliberate bug, along with a corrected version… See the full description on the dataset page: https://huggingface.co/datasets/creeperdatasets/python_debugging.texttext-generationn<1K0 likes80 downloads26d agoHugging Face19ArkhAngelLifeJiggy /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,161 TRAJECTORIES · 11,235 TRAINING ROWS · 11 MB PARQUET · 656 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K0 likes70 downloads2mo agoHugging Face20ansulev /fable-5-coding-and-debugging-traces Mirror: greghavens/fable-5-coding-and-debugging-traces Pinned snapshot / mirror of greghavens/fable-5-coding-and-debugging-traces, re-hosted for PROTISEC research reproducibility. Redistributed under the upstream license (cc-by-4.0) with attribution — all credit to the original author. Original author: greghavens Source dataset: greghavens/fable-5-coding-and-debugging-traces License: cc-by-4.0 Family: coding_traces Mode: stream Rows cached: 11487 Changes vs upstream: cached… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/fable-5-coding-and-debugging-traces.tabular10K<n<100K0 likes65 downloads2mo agoHugging Face2111-47 /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/11-47/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes63 downloads5d agoHugging Face22Precise-Debugging-Benchmarking /PDB-Single-Hard PDB-Single-Hard: Precise Debugging Benchmarking — hard single-line bug subset 📄 Paper  ·  💻 Code  ·  🌐 Project page  ·  🏆 Leaderboard PDB-Single-Hard is the hard single-line bug subset of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets: PDB-Single ·… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Hard.tabulartext-generation1K<n<10K0 likes61 downloads5mo agoHugging Face23stindardlogic /code-debugging-sft-50k Code Debugging SFT (50K) 50,000 ShareGPT-format conversations where the user presents buggy code and the assistant provides root-cause analysis and a corrected solution. Covers Python, JavaScript, Go, TypeScript, and SQL across 14 bug categories. Motivation Debugging is one of the most frequent developer tasks — and one of the hardest to train. Most coding datasets focus on writing code from scratch. This dataset trains models to: Identify the precise root cause… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/code-debugging-sft-50k.texttext-generation10K<n<100K0 likes61 downloads2mo agoHugging Face24tesraghavan /agent-traces-data-pipeline-debugging Agent Traces: data-pipeline-debugging Synthetic multi-agent workflow traces with LLM-enriched content for the data-pipeline-debugging domain. Part of the juliensimon/open-agent-traces collection — 10 datasets covering diverse domains and workflow patterns. What is this dataset? This dataset contains 2,033 events across 50 workflow runs, each representing a complete multi-agent execution trace. Every trace includes: Agent reasoning — chain-of-thought for each… See the full description on the dataset page: https://huggingface.co/datasets/tesraghavan/agent-traces-data-pipeline-debugging.tabular1K<n<10K0 likes60 downloads13d agoHugging Face25juliensimon /agent-traces-data-pipeline-debugging Agent Traces: data-pipeline-debugging Synthetic multi-agent workflow traces with LLM-enriched content for the data-pipeline-debugging domain. Part of the juliensimon/open-agent-traces collection — 10 datasets covering diverse domains and workflow patterns. What is this dataset? This dataset contains 2,033 events across 50 workflow runs, each representing a complete multi-agent execution trace. Every trace includes: Agent reasoning — chain-of-thought for each agent step… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/agent-traces-data-pipeline-debugging.tabular1K<n<10K0 likes54 downloads6mo agoHugging Face26Precise-Debugging-Benchmarking /PDB-Single PDB-Single: Precise Debugging Benchmarking — single-line bug subset 📄 Paper  ·  💻 Code  ·  🌐 Project page  ·  🏆 Leaderboard PDB-Single is the single-line bug subset of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets: PDB-Single-Hard · PDB-Multi… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single.tabulartext-generation1K<n<10K0 likes52 downloads5mo agoHugging Face2711-47 /fable-5-coding-and-debugging-traces-synthetic Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/11-47/fable-5-coding-and-debugging-traces-synthetic.tabulartext-generationn<1K0 likes52 downloads5d agoHugging Face28kalaiarasan27 /Code_Debugging_QA Code Debugging Q&A Dataset By dmeldrum6 A curated dataset of 1,073 question-and-answer pairs covering common debugging scenarios across Python, JavaScript, SQL, and Bash. Designed for fine-tuning and instruction-tuning language models on code debugging tasks. Dataset Summary Each pair presents a realistic bug symptom as a question and a structured answer containing: A buggy code block demonstrating the problem A corrected code block showing the fix A… See the full description on the dataset page: https://huggingface.co/datasets/kalaiarasan27/Code_Debugging_QA.text1K<n<10K0 likes39 downloads21d agoHugging Face29renjiepi /medium_20000-debugging_n100k1text1K<n<10K0 likes35 downloads8mo agoHugging Face30Precise-Debugging-Benchmarking /PDB-Multi PDB-Multi: Precise Debugging Benchmarking — multi-line bug subset (2–4 line blocks) 📄 Paper  ·  💻 Code  ·  🌐 Project page  ·  🏆 Leaderboard PDB-Multi is the multi-line bug subset (2–4 line blocks) of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets:… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Multi.tabulartext-generationn<1K0 likes34 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.