CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M3 likes4k downloads24d agoHugging Face02greghavens /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K65 likes1.6k downloads2mo agoHugging Face03greghavens /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K21 likes696 downloads2mo agoHugging Face04DSFFGFG456 /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,380 TRAJECTORIES · 12,490 TRAINING ROWS · 14 MB PARQUET · 663 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/DSFFGFG456/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K3 likes620 downloads2mo agoHugging Face05Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M0 likes517 downloads24d agoHugging Face06thoughtworks /agentic-coding-trajectories agentic-coding-trajectories A unified, tokenized corpus of 15,000 multi-turn agentic-coding sessions (618K turns, 41 turns/session avg) drawn from three publicly-released upstream datasets. Built for benchmarking LLM serving systems on realistic multi-turn coding-agent workloads. Why this exists Most LLM serving benchmarks use single-shot prompts. Real coding agents work in long multi-turn loops where each turn appends to a growing prompt. This corpus captures that shape… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/agentic-coding-trajectories.tabulartext-generation10K<n<100K1 likes464 downloads5mo agoHugging Face07jiajiale9 /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 697 TRAJECTORIES · 4,890 TRAINING ROWS · 3 MB PARQUET · 89 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/jiajiale9/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes273 downloads2mo agoHugging Face08Aratako /Synthetic-JP-EN-Coding-Dataset-801k Synthetic-JP-EN-Coding-Dataset-801k Magpieによって作成したコードSFTデータセットであるAratako/Synthetic-JP-EN-Coding-Dataset-Magpie-69kを元に、Evol-Instructのような手法を用いて複数のinstructionとresonseを生成し拡張して作成した、日英混合801262件のコードSFT用合成データセットです。 日本語: 173849件 英語: 627413件 元のinstructionの作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。 nvidia/Nemotron-4-340B-Instruct microsoft/Phi-3-medium-4k-instruct mistralai/Mixtral-8x22B-Instruct-v0.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-EN-Coding-Dataset-801k.tabulartext-generation100K<n<1M17 likes164 downloads2y agoHugging Face09ArkhAngelLifeJiggy /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes157 downloads2mo agoHugging Face10moehamid /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,374 TRAJECTORIES · 12,448 TRAINING ROWS · 14 MB PARQUET · 662 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/moehamid/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K0 likes139 downloads2mo agoHugging Face11greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes133 downloads2mo agoHugging Face12moehamid /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 601 TRAJECTORIES · 4,089 TRAINING ROWS · 3 MB PARQUET · 73 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/moehamid/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes130 downloads2mo agoHugging Face13siddharth0713 /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,161 TRAJECTORIES · 11,235 TRAINING ROWS · 11 MB PARQUET · 656 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/siddharth0713/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K0 likes116 downloads2mo agoHugging Face1411-47 /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/11-47/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes107 downloads6d agoHugging Face15Distillio /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/Distillio/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes100 downloads1mo agoHugging Face16witcheer /local-agentic-coding-bench-8gb-vram-2026-05 agentic coding benchmark: local LLMs on 8GB VRAM can local LLMs do agentic coding (multi-turn tool calling, file creation, debugging) on consumer hardware? this dataset captures real test results. hardware GPU: NVIDIA RTX 4060 Ti 8GB CPU: Intel i7-14700F RAM: 32 GB DDR5 OS: Windows 11 + WSL2 (Ubuntu) inference: llama-server (turboquant fork of llama.cpp) what was tested two agent frameworks: Hermes Agent (NousResearch): structured tool calling with… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/local-agentic-coding-bench-8gb-vram-2026-05.tabulartext-generationn<1K8 likes81 downloads4mo agoHugging Face17rashdan1 /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/rashdan1/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes77 downloads2mo agoHugging Face18squeezebits /synthesized-coding-assistant-dataset Synthesized Coding Assistant Dataset Overview Coding assistants are increasingly used for real-world software engineering workflows. However, there are relatively few datasets that closely resemble how such assistants operate in practice. Many existing coding datasets are based on single-turn or single-iteration tasks, where a model receives one coding request and directly produces an answer or patch. In contrast, practical coding assistants often work through… See the full description on the dataset page: https://huggingface.co/datasets/squeezebits/synthesized-coding-assistant-dataset.tabulartext-generationn<1K0 likes70 downloads4mo agoHugging Face19ArkhAngelLifeJiggy /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,161 TRAJECTORIES · 11,235 TRAINING ROWS · 11 MB PARQUET · 656 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K0 likes69 downloads2mo agoHugging Face2011-47 /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/11-47/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes64 downloads6d agoHugging Face2111-47 /fable-5-coding-and-debugging-traces-synthetic Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/11-47/fable-5-coding-and-debugging-traces-synthetic.tabulartext-generationn<1K0 likes55 downloads6d agoHugging Face22hbudhi36 /synthetic-coding-tutor Synthetic Coding Tutor Conversations Multi-turn student–tutor debugging conversations generated by a LangGraph pipeline of autonomous LLM agents. Each candidate solution is executed against real pytest test cases, and every finished conversation is graded 0–10 by an LLM judge on persona fidelity, tutor responsiveness, and dialog flow. Summary Total conversations: 2495 Training-ready (gold + silver): 1241 Quality buckets: gold=824, silver=417, bronze=1254… See the full description on the dataset page: https://huggingface.co/datasets/hbudhi36/synthetic-coding-tutor.tabulartext-generation1K<n<10K0 likes50 downloads14d agoHugging Face23nayohan /gpt4o-coding-eval-by-gemini1_5flash-koTranslated llama-duo/gpt4o-coding-eval-by-gemini1_5flash using nayohan/llama3-instrucTrans-enko-8b. This dataset is a raw translated dataset and contains repetitive sentences generated by the model, so it needs to be filtered. tabulartext-generationn<1K0 likes41 downloads2y agoHugging Face24Self-Improving-Coding-Agents /SI2CA-Training-TrajectoriesDataset Card for SI2CA-Training-Trajectories [🌐 Website] • [🤗 Dataset] • [📜 Paper] • [🐱 GitHub] 💡 Introduction This dataset consists of 32,340 coding-agent trajectories generated by Qwen3.5-122B-A10B on the same 10,780 executable Python SWE tasks under the three trajectory-curation settings of Section 4.4 of the paper: standard sampling, full self-judgement, and an efficient discovered strategy found by the recursive self-improvement framework. Each task is… See the full description on the dataset page: https://huggingface.co/datasets/Self-Improving-Coding-Agents/SI2CA-Training-Trajectories.tabulartext-generation10K<n<100K0 likes40 downloads1d agoHugging Face25davidkling /hf-coding-tools-dashboard-v2 HuggingFace AI Coding Tools Dashboard (Enhanced) Enhanced benchmark data from the HuggingFace AI Dashboard — includes query metadata (query_set, intent), run metadata (run_name, run_date), and freshness flags for stale references. This is the v2 enhanced dataset. The original dataset is at davidkling/hf-coding-tools-dashboard. Dataset Structure Split Description Rows results Enhanced results with query/run metadata and freshness flags 9146 queries… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard-v2.tabulartext-generation1K<n<10K1 likes38 downloads5mo agoHugging Face26davidkling /hf-coding-tools-dashboard-all HuggingFace AI Coding Tools Dashboard Benchmark data from the HuggingFace AI Dashboard — tracking how AI coding tools (Claude Code, Codex, Copilot, Cursor) recommend HuggingFace products across 32 developer categories. Dataset Structure Split Description Rows results Full benchmark results with LLM responses, cost, tokens, latency, and product detection 9603 queries Benchmark query definitions across 32 categories 404 runs Run metadata and tool/model… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard-all.tabulartext-generation10K<n<100K0 likes38 downloads4mo agoHugging Face27davidkling /hf-coding-tools-dashboard HuggingFace AI Coding Tools Dashboard Benchmark data from the HuggingFace AI Dashboard — tracking how AI coding tools (Claude Code, Codex, Copilot, Cursor) recommend HuggingFace products across 32 developer categories. Dataset Structure Split Description Rows results Full benchmark results with LLM responses, cost, tokens, latency, and product detection 9146 queries Benchmark query definitions across 32 categories 404 runs Run metadata and tool/model… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard.tabulartext-generation1K<n<10K0 likes36 downloads5mo agoHugging Face28davidkling /hf-coding-tools-dashboard-run-april12 HuggingFace AI Coding Tools Dashboard Benchmark data from the HuggingFace AI Dashboard — tracking how AI coding tools (Claude Code, Codex, Copilot, Cursor) recommend HuggingFace products across 32 developer categories. Dataset Structure Split Description Rows results Full benchmark results with LLM responses, cost, tokens, latency, and product detection 8875 queries Benchmark query definitions across 32 categories 263 runs Run metadata and tool/model… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard-run-april12.tabulartext-generation1K<n<10K0 likes34 downloads4mo agoHugging Face29davidkling /hf-coding-tools-dashboard-builder HuggingFace AI Coding Tools Dashboard Benchmark data from the HuggingFace AI Dashboard — tracking how AI coding tools (Claude Code, Codex, Copilot, Cursor) recommend HuggingFace products across 32 developer categories. Dataset Structure Split Description Rows results Full benchmark results with LLM responses, cost, tokens, latency, and product detection 581 queries Benchmark query definitions across 32 categories 120 runs Run metadata and tool/model… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard-builder.tabulartext-generationn<1K0 likes31 downloads4mo agoHugging Face30InfoBayAI /DSA-Coding-Problems-and-Solutions-Datasetgated Dataset Description This dataset is a large-scale collection of Data Structures and Algorithms (DSA) code, containing 12,385 code files with 3.86 million lines of code and 25.01 million lexical tokens, designed to support the development of advanced code generation models, programming assistants, software engineering AI systems, and code intelligence applications. It consists of real-world DSA implementations covering a wide range of algorithms, data structures, problem-solving… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/DSA-Coding-Problems-and-Solutions-Dataset.tabulartext-generationn<1K0 likes23 downloads7d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.