CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Open-SWE-Traces Open-SWE-Traces: Advancing Distillation for Software Engineering Agents 🚨 What's New [09/26] Release v1.2: Added new agent trajectories generated by Qwen3.8-27B for mini-swe-agent. Trajectories for OpenCode and Claude Code harnesses will be released soon. [08/26] Release v1.1: Added new agent trajectories generated by DeepSeek-V4-Flash and Qwen3.6-27B across OpenHands, SWE-agent, and mini-swe-agent harnesses. [06/21] Release v1.0: Released 207k agent… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Open-SWE-Traces.text100K<n<1M131 likes32k downloads3d agoHugging Face02semianalysisai /cc-traces-weka-062126 semianalysisai/cc-traces-weka-062126 WekaTrace corpus derived from SemiAnalysis Claude Code proxy traces. Built 2026-06-21 17:48:24 UTC via utils/agentic/build_weka_hf_dataset.py. Filters Trace version: exactly v7 min Anthropic requests per session: 20 Claude Code CLI ≥ 2.1.139 (every row) peak concurrent sub-agent groups ≤ 10 Non-image rows only (image content excluded at source) Classifier calls excluded (max_tokens<=64 AND no tools → SUGGESTION MODE, title-gen… See the full description on the dataset page: https://huggingface.co/datasets/semianalysisai/cc-traces-weka-062126.texttext-generationn<1K10 likes28k downloads3mo agoHugging Face03semianalysisai /cc-traces-weka-062126-256k semianalysisai/cc-traces-weka-062126-256k WekaTrace corpus derived from SemiAnalysis Claude Code proxy traces. Built 2026-06-21 17:49:45 UTC via utils/agentic/build_weka_hf_dataset.py. Derived from semianalysisai/cc-traces-weka-062126 by applying the 256k per-request cap and preserving the surviving requests' relative timestamps. Filters Trace version: exactly v7 min Anthropic requests per session: 20 Claude Code CLI ≥ 2.1.139 (every row) peak concurrent… See the full description on the dataset page: https://huggingface.co/datasets/semianalysisai/cc-traces-weka-062126-256k.texttext-generationn<1K5 likes20k downloads3mo agoHugging Face04cot-leaderboard /cot-eval-traces-2.0text1M<n<10M9 likes17k downloads2y agoHugging Face05Infatoshi /kernelbench-mega-traces KernelBench-Mega agent traces Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells). Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score. 23 agent traces · live leaderboard: https://kernelbench.com/mega Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.tabularn<1K17 likes6.7k downloads4h agoHugging Face06MaxDevv /real-pi-coding-agent-traces-sessions Real Pi Coding Agent Traces Sessions An aggregated dataset of real human–AI coding agent sessions, collected from 21 independently published Hugging Face datasets and hand-filtered to exclude synthetic or AI-generated content. Every session is an unedited (but redacted) trace of a real person using pi — an open-source AI coding agent harness — to build, debug, and ship real open-source software. Real prompts, real tool calls, real errors, real backtracking. Why this… See the full description on the dataset page: https://huggingface.co/datasets/MaxDevv/real-pi-coding-agent-traces-sessions.text-generation1K<n<10K4 likes6.2k downloads2mo agoHugging Face07masterpieceexternal /gpt-oss-20b-moe-expert-power-traces-320k GPT-OSS-20B MoE Expert Power Traces (320k, ChipWhisperer) This dataset contains analog power traces captured with a ChipWhisperer Husky while running forced single-expert MoE computations derived from openai/gpt-oss-20b on an NVIDIA H100. What is recorded Each trace corresponds to one capture trial where: A fixed expert id is selected (expert_00 ... expert_31). A random hidden-state tensor is generated once per trial. The selected expert computation is executed… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-moe-expert-power-traces-320k.audio-classification100K<n<1M0 likes5.8k downloads4mo agoHugging Face08Infatoshi /kernelbench-hard-traces KernelBench-Hard agent traces Frontier coding agents writing optimized CUDA/Triton kernels (FP8 GEMM, paged attention, MoE, W4A16, KDA, Top-k) on RTX PRO 6000 Blackwell, H100 PCIe, and B200; roofline-graded. Each .jsonl file is one agent run in Claude-Code session format, viewable with the Hugging Face Agent Trace viewer (Data Studio → open a row). Filename = run id. Live leaderboard: https://kernelbench.com/hard Secrets redacted. Full reasoning for open-provider routes… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-hard-traces.tabulartext-generationn<1K16 likes5.3k downloads7d agoHugging Face09imo2026-challenge /chankhavu-imo-reasoning-traces0 likes4.6k downloads2mo agoHugging Face10dacorvo /funes-nvidia-Open-SWE-Traces Funes recall store — NVIDIA Open-SWE-Traces (resolved) A funes recall store built by indexing the resolved==1 trajectories of nvidia/Open-SWE-Traces (65244 sessions, across both harnesses — SWE-agent and OpenHands — and both models, Minimax-M2.5 and Qwen3.5-122B). What this is This is not a raw trace dataset — it is a pre-built funes index: the source trajectories chunked into content blocks and embedded, stored as a Lance table (chunks.lance). Source… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-nvidia-Open-SWE-Traces.10M<n<100M0 likes4.5k downloads2mo agoHugging Face11viktor-shcherb /jobseek-agent-traces Jobseek Agent Traces Claude Code agent session traces from jobseek — a job posting monitor for company career pages. Each trace captures a complete agent workflow session: company discovery, board configuration, monitor/scraper selection, and quality validation. These are raw session transcripts, not tabular data — use the trace viewer to explore them. Structure traces/ {company-slug}/ {date}.jsonl # One trace per session (header + records) Each .jsonl… See the full description on the dataset page: https://huggingface.co/datasets/viktor-shcherb/jobseek-agent-traces.1K<n<10K4 likes4.4k downloads0m agoHugging Face12nlile /misc-merged-claude-code-traces-v1 MISC Unification of Public Claude Code Traces A unified dataset of 32,133 deduplicated Claude API conversation traces focused on software engineering and code generation tasks. This dataset merges and normalizes traces from 10 different source datasets into a single, consistent format. Dataset Description This dataset contains real Claude API interaction traces capturing software engineering workflows including: Code generation and modification Bug fixing and debugging… See the full description on the dataset page: https://huggingface.co/datasets/nlile/misc-merged-claude-code-traces-v1.text10K<n<100K19 likes3.7k downloads9mo agoHugging Face13agent-evals /hal_traces8 likes3.2k downloads8mo agoHugging Face14TeichAI /Ox-Alpha-Pi-TracesThis dataset was generated using teich by TeichAI Ox-Alpha Pi Agent Coding Traces This directory contains raw agent trace files generated by teich. JSONL files: 2247 Model metadata: stealth/ox-alpha Domains and prompt distribution Topic Traces Games & simulation (headless) 196 Frontend & Node-testable web 159 Health & medicine informatics 139 ML & scientific computing (CPU) 123 Data analysis & reporting 122 Computational biology & chemistry… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Ox-Alpha-Pi-Traces.text-generation8 likes3.1k downloads28d agoHugging Face15jonathanyin /aime_1983_2023_deepseek-r1_traces_16384tabularn<1K0 likes2.9k downloads1y agoHugging Face16fol-traces /fol-traces citation @misc{lee2025foltraces, title={FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale}, author={Lee, Isabelle and Liaw, Sarah and Yogatama, Dani}, year={2025}, eprint={2505.14932}, archivePrefix={arXiv}, primaryClass={cs.AI}, url={https://arxiv.org/abs/2505.14932} } text-generation1B<n<10B3 likes2.6k downloads3mo agoHugging Face17choucsan /mimo-claude-code-traces-1k MIMO Claude Code Traces MIMO Claude Code Traces is a collection of coding-agent trajectories in a Claude Code-style environment. Each record contains a user coding task, the full multi-turn message trace, available tool schemas, assistant reasoning fields, tool calls, tool outputs, and metadata such as model name, category, duration, cost, token usage, and whether the trace used tools. The traces were generated with mimo-v2.5-pro, MiMo's most capable model at the time of… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/mimo-claude-code-traces-1k.tabulartext-generation1K<n<10K11 likes2.6k downloads2mo agoHugging Face18experiential-labs /wmo-terminal-tasks-traces terminal-tasks — real agent-environment traces Computer-use agent runs in real terminal containers: bash commands and their true outputs from live task environments. Every trace is a REAL run: an LLM agent stepping against the actual benchmark environment, with each transition (tool call → true environment observation) recorded as OpenTelemetry GenAI spans (traces.otel.jsonl, one span per line). Captured by world-model-harness's environment-capture package, which also holds the… See the full description on the dataset page: https://huggingface.co/datasets/experiential-labs/wmo-terminal-tasks-traces.2 likes2.5k downloads2mo agoHugging Face19dungnv /qwen36-27b-length-traces Qwen3.6-27B generation-length prediction: heads, calibrations and workloads Artifacts for conformal length-aware LLM scheduling on Qwen/Qwen3.6-27B — predicting a request's remaining generation length from a hidden layer during decoding, wrapping it in a split-conformal interval, and scheduling with SRPT inside vLLM. Extends TRAIL (Don't Stop Me Now, ICLR'25) to a hybrid-attention reasoning model. This repo contains the derived artifacts, not the raw activations. The 3250… See the full description on the dataset page: https://huggingface.co/datasets/dungnv/qwen36-27b-length-traces.text-generation0 likes2.4k downloads1mo agoHugging Face20devichand /schedulerlens-traces1 likes2.3k downloads2mo agoHugging Face21Exgentic /agent-llm-traces-v2 Exgentic Agent LLM Traces v2 — Agent Chat Only OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it. This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.tabulartext-generation10K<n<100K0 likes2.3k downloads3mo agoHugging Face22smolagents /post-train-bench-traces PostTrainBench Sessions by Benchmark Derived from akseljoonas/posttrainbench-sessions on 2026-04-20. This dataset exports each source row as one viewer-compatible JSONL trace and groups traces by benchmark. Layout benchmarks.json: benchmark catalog and counts benchmarks/<benchmark>/index.json: metadata index for one benchmark benchmarks/<benchmark>/<job_id>.jsonl: one converted session trace per source row Benchmarks Benchmark Sessions aime2025 19… See the full description on the dataset page: https://huggingface.co/datasets/smolagents/post-train-bench-traces.0 likes2.1k downloads5mo agoHugging Face23amanutej /trustworthy-biology-agents-traces Trustworthy Biology Agents — Run Traces Raw execution traces from 1,329 agent runs across three coding agents on three biology benchmarks — BiomniBench-DA, BixBench, and CompBioBench. This is the scrubbed trace bundle for the study in manu-tej/ai-scientists; the write-up lives in that repo's RESULTS.md. The motivating question is not only whether an agent reaches the right answer, but whether it behaves like a trustworthy analyst when the task is ambiguous, under-specified, or… See the full description on the dataset page: https://huggingface.co/datasets/amanutej/trustworthy-biology-agents-traces.tabular1K<n<10K0 likes2.1k downloads2mo agoHugging Face24trace-commons /agent-traces Trace Commons — Agent Traces Trace Commons is one open, public dataset of coding-agent sessions — the back-and-forth between a developer and an AI coding agent, including prompts, model responses, tool calls, and command output — contributed voluntarily as an open resource for studying, evaluating, and building on how these agents actually work. Every trace here was donated only from a public, open-source repository, was anonymized on the contributor's own machine before upload… See the full description on the dataset page: https://huggingface.co/datasets/trace-commons/agent-traces.tabulartext-generationn<1K35 likes2k downloads3mo agoHugging Face25RESMP-DEV /Fable-GPT-5.5-Distillation-Traces Agent Traces Curated 2026 (v3 Merged) A unified distillation corpus of 9,057,143 records spanning agentic coding traces, math/code/science reasoning, tool-use trajectories, and preference data. 8,876,012 train + 181,131 eval, stratified by source. What this is This is the v3 merged corpus that supersedes both v1 and v2 of this dataset. It combines five major source groups through a unified normalization pipeline: Original v2 RESMP-DEV (de-fragmented, re-deduped):… See the full description on the dataset page: https://huggingface.co/datasets/RESMP-DEV/Fable-GPT-5.5-Distillation-Traces.texttext-generation1M<n<10M10 likes2k downloads3mo agoHugging Face26uynitsuj /paper-sim-n512-traces paper-sim-n512-traces Rollout traces for the n=512 simulated bottle-in-bin evaluation reported in Table I of the WARP-RM CoRL rebuttal. Two arms, 512 paired scenes each: directory prefix arm bottles/scene throughput all-6 fullhz_vanilla_sh00..15 Vanilla BC (100% of data) 3.885 237/hr 9.4% fullhz_paperwarp512_sh00..15 WARP-BC (31.5% kept) 4.533 290/hr 25.0% 512 traces per arm = 16 shards x 32 worlds. Each qpos_trace_NNN_sSEED.npz holds the full qpos trajectory… See the full description on the dataset page: https://huggingface.co/datasets/uynitsuj/paper-sim-n512-traces.robotics0 likes1.9k downloads1mo agoHugging Face27Lottolabs /terminal-bench-2.1-qwen3.8-27b-traces Terminal-Bench 2.1 traces: Qwen3.8-27B-GPTQ-4bit, xhigh / medium / low / off Complete agent trajectories, verifier output, timing and token usage for all 89 Terminal-Bench 2.1 tasks run locally with btbtyler09/Qwen3.8-27B-GPTQ-4bit on 2× RTX 3090, plus the adaptive fallback reruns at lower reasoning effort. Headline result: 62/89 (69.66%) at xhigh in a single clean pass. Cumulative best-of across xhigh → medium → low → off fallbacks: 70/89 (78.65%). The second number is not a… See the full description on the dataset page: https://huggingface.co/datasets/Lottolabs/terminal-bench-2.1-qwen3.8-27b-traces.text1K<n<10K1 likes1.9k downloads18d agoHugging Face28albertoRodriguez97 /history-anchor-100-traces History Anchor 100 — Model Trajectories *Per-(model × condition × scenario set × seed) raw outputs from the paper "History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions".* This dataset contains the full set of model decisions that back every figure and table in the paper. Use it to: audit a single model's behaviour scenario-by-scenario, recompute headline metrics without re-running the (paid) API sweeps, mine reasoning_content traces from models that expose… See the full description on the dataset page: https://huggingface.co/datasets/albertoRodriguez97/history-anchor-100-traces.text-generation10K<n<100K0 likes1.7k downloads4mo agoHugging Face29mlfoundations-dev /terminal-bench-traces-localtext1K<n<10K0 likes1.7k downloads1y agoHugging Face30greghavens /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K65 likes1.6k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.