datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
debugtest3
DebugTestkimi-k3-coding-and-debugging-traces
Kimi K3 Coding, Tool Use & Instruction Following Traces
582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables
below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/kimi-k3-coding-and-debugging-traces.super-debug-v3
super-debug-v3
super-debug-v3 is a synthetic dataset of grounded software-debugging trajectories generated with hen, a stateful long-horizon AI coding agent for C/C++ projects.
This is the third version of super-debug. Compared with super-debug-v2, which focused on SimpleC/compiler debugging runs, v3 includes synthesized data across the newer hen/Projects project set:
clcalc
math3d
mini2d_tilegame
ocr8
poseblend
rigid2d
sgps
simplec
tinyvm
The default config is the… See the full description on the dataset page: https://huggingface.co/datasets/georvn7/super-debug-v3.DebugBench
Dataset Summary
DebugBench is a Large Language Model (LLM) debugging benchmark introduced in the paper DebugBench: Evaluating Debugging Capability of Large Language Models. We collect code snippets from the LeetCode community and implant bugs into source data with GPT-4. The project is also open-sourced as a GitHub repository.
It consists of 4,253 instances.
It covers four major bug categories and 18 minor types.
It includes C++, Java, and Python instances.
It contains three… See the full description on the dataset page: https://huggingface.co/datasets/Rtian/DebugBench.super-debug-v2
super-debug-v2
super-debug-v2 is a synthetic dataset of grounded software-debugging trajectories generated with hen, a stateful long-horizon AI coding agent for C/C++ projects.
This is the second version of super-debug. Compared with super-debug-v1, this release is generated from three full-suite debugging runs. The previous release kept trajectories that passed only the first three validation steps; this version keeps trajectories from runs that pass the full hen/SimpleC/tests… See the full description on the dataset page: https://huggingface.co/datasets/georvn7/super-debug-v2.glm-5.2-coding-and-debugging-traces
GLM 5.2 Agent Traces
207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from GLM 5.2 (glm-5.2). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task domain.
This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/glm-5.2-coding-and-debugging-traces.fable-5-coding-and-debugging-traces
Claude Fable 5 Agent Traces
2,380 TRAJECTORIES · 12,490 TRAINING ROWS · 14 MB PARQUET · 663 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task… See the full description on the dataset page: https://huggingface.co/datasets/DSFFGFG456/fable-5-coding-and-debugging-traces.cached-activationsrag-retrieval-debug-trajectories
Rag Retrieval Debug Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/rag-retrieval-debug-trajectories.LoVR-benchmarkThis repository hosts the LoVR benchmark dataset, designed for research on long video–text retrieval.The original video data is sourced from LongVideoBench. We sincerely thank the authors for their outstanding work.
Building upon LongVideoBench, we curate high-quality captions and semantic annotations tailored for long-form video understanding and retrieval tasks.
The corresponding GitHub repository for this project is available at:👉 https://github.com/TechNomad-ds/LoVR-benchmark
Below we… See the full description on the dataset page: https://huggingface.co/datasets/debugger123/LoVR-benchmark.kimi-k3-coding-and-debugging-traces
Kimi K3 Coding, Tool Use & Instruction Following Traces
697 TRAJECTORIES · 4,890 TRAINING ROWS · 3 MB PARQUET · 89 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables
below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/jiajiale9/kimi-k3-coding-and-debugging-traces.cua_debugger_traj
CUA Debugger Trajectories
204 failed computer-use agent (CUA) trajectories on OSWorld, each with a human root-cause annotation.
Three agents were run on OSWorld (Ubuntu desktop, screenshot-only observation, pyautogui execution at 1920×1080). Every trajectory in this dataset is a failure (no task reached evaluator score 1.0). For each trajectory, a human annotator identified the root error step — the earliest step responsible for the failure — and labeled it with an… See the full description on the dataset page: https://huggingface.co/datasets/CyT1ng/cua_debugger_traj.generations-21-DEBUG-qwen3-8b-simnpo-gentle-igm-10b-target-100-localtrain-checkpoint-1generations-17-DEBUG-qwen3-8b-simnpo-gentle-baseline-target-100-localtrain-checkpoint-1generations-18-DEBUG-llama-3_1-8b-simnpo-gentle-bm25-10b-target-100-localtrain-checkpoint-1observability-debug-trajectories
Observability Debug Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/observability-debug-trajectories.glm-5.2-coding-and-debugging-traces
GLM 5.2 Agent Traces
207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from GLM 5.2 (glm-5.2). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task domain.
This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/glm-5.2-coding-and-debugging-traces.1016_SWElite_debugdebug_MMMU_mcq_to_remove
Dataset Card for "debug_MMMU_mcq_to_remove"
More Information needed
generations-checkpoint-134-debug-checkpoint-134-llamafable-5-coding-and-debugging-traces
Claude Fable 5 Agent Traces
2,374 TRAJECTORIES · 12,448 TRAINING ROWS · 14 MB PARQUET · 662 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task… See the full description on the dataset page: https://huggingface.co/datasets/moehamid/fable-5-coding-and-debugging-traces.nemotron-terminal-debugging
nemotron-terminal-debugging
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "debugging". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-debugging.fable-5-coding-and-debugging-traces-synthetic-corrections
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.DataCat1k
CatDataset1k
Dataset of 1000 images of cats (domestic cats, Felis catus) for training models,
experiments and fine-tuning (image generation, classification, etc.).
Query: cat
Caption / label for every image: cat
Files: cat_0000.jpg ... cat_0999.jpg (JPEG)
Sources: Wikimedia Commons + Flickr (via Openverse), open licenses
How to download / use
1. Load directly with the datasets library (recommended)
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/debugdll/DataCat1k.kimi-k3-coding-and-debugging-traces
Kimi K3 Coding, Tool Use & Instruction Following Traces
601 TRAJECTORIES · 4,089 TRAINING ROWS · 3 MB PARQUET · 73 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables
below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/moehamid/kimi-k3-coding-and-debugging-traces.global-news-radio-debug
Global News Radio Dataset (1 hour per station)
Every news radio station from the Radio Browser API, recorded for 1 hour each.
Attempted
3037
Successful
2553
Failed
484
Total audio
21 hours
Parquet shards
256
Size
0.6 GB
Format
MP3 16kHz mono 64kbps
Usage
from datasets import load_dataset
ds = load_dataset("NathanRoll/global-news-radio-debug", streaming=True)
for sample in ds["train"]:
print(sample["station_name"], sample["language"]… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/global-news-radio-debug.generations-16-DEBUG-llama-3_1-8b-simnpo-gentle-baseline-target-100-localtrain-checkpoint-1whisper-non-verbal-debugfable-5-coding-and-debugging-traces
Claude Fable 5 Agent Traces
2,161 TRAJECTORIES · 11,235 TRAINING ROWS · 11 MB PARQUET · 656 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task… See the full description on the dataset page: https://huggingface.co/datasets/siddharth0713/fable-5-coding-and-debugging-traces.
